Premise: Data are used to answer questions and provide insight into problem-solving. Being able to meaningfully interpret of data is one of the most important skills that any scientist can have. Those skills are derived from experience and an understanding of the strength and weakness of raw data and different forms of analyses.
I. Hypothesis testing
A. Essential concepts
1. Scientists seek to form generalizations from individual observations. That is inductive logic.
2. Generalizations can either be valid or invalid. We must determine whether a given generalization is valid.
3. Generalizations based on relatively scanty evidence are called hypotheses. Since they are so tentative, hypotheses must be tested to determine their validity.
B. Philosophy of hypothesis testing
1. Whole process is similar to that of a judge seeking to determine whether a particular person is guilty or innocent. However, the judge can only make a best guess based on the evidence. Thus the judge decides to acquit or convict.
2. We want to acquit innocent people and convict guilty ones. However, sometimes mistakes are made.
3. Researchers "judge" hypothesis. Based on the evidence, we decide to accept or reject.
C. Hypothesis testing - nuts and bolts
1. Involves examining a new set of specific cases to determine whether they fit the generalization. This is deductive logic.
2. Often the specific cases are not identical to each other. Thus we must employ a sampling approach and generate population statistics.
3. Often the observed results differ to some degree from the expected. We then must determine whether the deviation is sufficiently large to warrant rejection of the hypothesis. That is done by statistical tests.
D. Statistical tests
1. Have a wide variety, including t-tests, chi square, analysis or variance, Mann-Whitney U-test, etc.
a. Each test has its own purpose
b. Each has its own assumptions
2. Involve collecting data and applying them to a formula
3. Generally get a test statistic. Compare to a tabular value. If the number is greater than the value we decide to reject or accept the hypothesis, depending on the test.
4. Tests are often based on P values. This is the proportion of time that we are willing to be wrong if we reject a hypothesis. Thus, a P value of 0.05 means that we are willing to be wrong 5% of the time if we reject a hypothesis using a particular statistical value.
E. Show example - t-test.
1. Used to compare population means.
2. Assumption is that we have large populations and each is normally distributed
3. Set up null hypothesis, in which we state that there is no difference between means
4. Calculate means and variances. Determine population size of each.
5. Calculate t statistic which is given as:
6. Note that if difference between means is small, t will be a small number. If difference is large, t will be large.
7. We want to reject the null hypothesis (that there is no difference between the means) when t is large.
8. We compare the calculated t value to a tabular value to assess whether the difference is significant.
II. Pitfalls to avoid when analyzing data
1. Make sure sample size is adequate.
A. Small samples can give misleading results. For example, two populations can have vastly different means, but if they are based on few samples the difference might not be significant.
B. Some people even draw inference from a single observation - testimonials can be misleading.
2. Trying to make too much out of small differences (NBA syndrome)
3. Sometimes statistically significant trends might not be meaningful in a real-world context (drug that can extend one's life by five minutes).
4. Sometimes analysis used is not appropriate to data - must use insight
5. Figures can be misleading when axes do not drop to zero, or when there are breaks.
Return
to Homepage for FRF 101O
This page posted and maintained by Kenneth M. Klemow, Ph.D., Biology Department, Wilkes University, Wilkes-Barre, PA 18766. (570) 408-4758, kklemow@wilkes.edu.