SOLUTIONS MANUAL
, The Analysis of Biological Data - Whitlock and Schluter
Solutions to assignment problems
Chapter 1
10. (a) Discrete (b) Continuous (c) Continuous (d) Discrete (e) Continuous
11. Observational study. The researcher has no control over which women have miscarriages and
which lose their fetus from other causes.
12. (a) numerical, discrete (the variable, if not the partners)
(b) numerical , continuous
(c) categorical, ordinal
(d) numerical, continuous
(e) categorical, ordinal
(f) numerical, continuous
(g) categorical, nominal
(h) numerical, discrete
(i) categorical, nominal
(j). numerical, continuous
13. (a) Observational study: the individual fish were not assigned to subspecies by the researcher.
(b) Subspecies of fish and wavelength of maximum sensitivity.
(c) The explanatory variable is the subspecies, the response variable is the wavelength of
maximum retina sensitivity.
14. (a) No. The 500 households selected to receive the survey might be a random sample, but the
low completion rate (< 20%) makes the sample a volunteer sample.
(b) Volunteer bias Those who volunteer to respond to a survey on recycling might have different
opinions of the program than those who did not respond.
15. (a) Omitting cell phones could bias a sample. If younger individuals are more likely to use a cell
phone, omitting cell phones would bias the sample towards older individuals. (b) Equal chance
of being selected.
16. (a) The population of interest is coastal Californian population of piñon pine trees.
(b) A single plot was randomly sampled, but trees were not randomly sampled. The multiple
trees within the same plot might not be independent, if they are related, of similar age, or share
the same environment.
17. The 60 samples are not a random sample. The 6 dives measured on each bird are not
independent. The six dive results measured from each bird are likely to be more similar than dive
results obtained from six different birds sampled randomly from the population.
,Chapter 2
14. (a) Between 12 and 13 mm.
(b) Approximately 50% of the finches are at the modal beak width.
(c) Changing the widths of intervals or “bins” of the histogram can alter its shape. Draw several
histograms with the data, using wider and narrower intervals, is needed to determine whether a
second peak is present. (d) Bimodal.
15. (a) Touching the first segment of the hind leg led to the greatest response. Touching the thorax or
distal portions of any of the legs resulted in the lowest response.
(b) Map.
16. (a) Frequency table.
(b) A single variable (number of convictions).
(c) 21.
(d) 265 of 395 (the fraction 0.67) had no convictions.
(e)
Histogram − it is the easiest way to visualize the frequency distribution for a numerical variable.
(The cumulative frequency distribution is also an appropriate graph).
(f) Skewed (right) and unimodal (mode is 0 convictions). There are no outliers.
(g) The sample was six schools near the research office ⎯ not a random sample of British boys
or any other population.
17. (a) This is a contingency table.
(b)
(c) Categorical, ordered. Groups should be arranged by increasing income.
(d) The relative frequency of conviction decrease as available income increases.
, (e) The mosaic plot made it easier to see the pattern. Whereas the table gives the frequencies, the
graph visualizes the association between the variables.
18. (a)
(b) Histogram: it visualizes the frequencies of each spermatophore mass interval very clearly.
(c) The main part of the distribution is fairly symmetric with a mode of 0.06−0.07. There is an
extreme measurement at large spermatophore mass.
(d) Outlier.
19. (a) Both variables are continuous numeric variables.
(b) Scatter plot.
(c) The relationship is positive but non-linear. As temperature increases the fusion frequency
increases.
(d) The 20 measurements are not a random sample because each fish was measured several times
and the multiple measurements were all combined.
20. (a) Line graph.
(b) The steepness of each segment tells us the net increase in the number of endangered species
added in a given year (it is not exactly the total number added, because some species might
have been removed from the list in a given year).
(c) The net number of endangered species has been increasing steadily over time, but has tapered
off toward the most recent dates.
21. (a) Histogram.
(b) The bars of a histogram should not have gaps between them. (A lesser problem is that it is
not clear what the ticks on the x-axis refer to.)
(c) The variation in protein similarity is the most interesting feature: some proteins are nearly
identical between humans and puffer fish, while others are nearly completely dissimilar.
(d) Skewed left.
(e) The mode is 70% similarity (presumably the interval the number “70” represents is
67.5−72.5).
22. (a) Cumulative frequency distribution.
(b) The y-axis indicates the quantile of the variable indicated on the x-axis (annual percent
change in human population). The quantile is the fraction of observations less than or equal to
the value on the x-axis.
(c) Approximately 10% of the countries had negative change in population size.
(d) The 0.10 quantile is approximately 0 growth, the 0.50 quantile is 1.5% growth, and the 0.90
quantile is 3% growth.
, (e) The 60th percentile is approximately 1.75% growth.
23. (a) Scatter plot.
(b) Number of fruits previously produced, because we wish to use it to predict photosynthetic
capacity.
(c) Negative association: photosynthetic capacity reduced in trees that produced many fruits
previously.
24. (a) Grouped histograms. Explanatory variable: genotype at PTC gene. Response variable: Taste
sensitivity score. Genotype is categorical variable, taste sensitivity is numerical. (b). Scatter plot.
Explanatory variable: migratory activity of parents. Response variable: migratory activity of
offspring. Both variables are numerical.
(c) Grouped cumulative frequency distributions. Explanatory variable: year of study. Response
variable: density of fine roots. Root density is numerical. While year is a numerical variable,
strictly speaking, it is used as a categorical variable in this figure to define the three groups of
measurements.
(c) Grouped bar graph. Explanatory variable: HIV status. Response variable: needle sharing.
Both variables are categorical.
25. (a) Percentage of adults with BMI greater than 25 increased steadily from 1995 to 2002 then it
dropped slightly and became steady after 2002.
(b) While cute, the figure does not help the eye visualize the association between year and the
percentage of adults with BMI greater than 25.
(c) Line graph.
26. (a) Contingency table.
No sneaker One sneaker Two or more sneakers Total
Eggs eaten 61 18 16 95
No eggs eaten 389 17 4 410
Total 450 35 20 505
, (b) Mosaic plot. (A grouped bar plot might also be effective.)
Chapter 3
10. (a) 5.5 (in log10 units).
(b) 0.26 recruits (in log10 units).
(c) 39/39 = 1.0 (100%).
11. (a) Box plot: (Grouped histogram or grouped cumulative frequency distribution are also valid.)
(b) V1a enhanced group has a higher mean (86%) than control (58%).
(c) Control group has the higher standard deviation (29.8%) than V1a enhanced group (12.9%).
12. (a) Histogram shows a sharply right-skewed frequency distribution of ages, with the mode at a
young age. There might be a second, low peak at intermediate ages.
, (b) Median appears to be between 0 and 5 million years ago (mya), whereas mean is between 5
and 10 mya. The mean is greater than the median because the distribution is right-skewed: the
large values influence the mean more than the median.
(c) Mean (8.66 mya) is indeed greater than the median (3.51 mya).
(d) First quartile: 1.105 mya; third quartile: 17.61 mya; interquartile range: 16.50 mya.
(e) Box plot:
13. (a) Median: 8.0 (the value of the 64th sorted observation).
(b) First quartile: 3 prey species. Third quartile: 17 prey species. Interquartile range: 14 prey
species.
(c) No, because we don't have the numbers in the "more than 20" class..
14. (a) This is a histogram.
(b) Mean: approximately 1000 yards/minute. The frequency distribution is fairly symmetric, so
the mean should lie near the middle.
(c) Median: approximately 900 yards/minute. The frequency distribution is fairly symmetric, so
the median should lie near the middle, close to the mean.
(d) Mode: 1000−1100 yards/minute (the most frequently occurring interval in a frequency
distribution)
(e) Standard deviation (s): approximately 200 yards/minute. Based on the fact that if the
distribution is roughly bell-shaped (normal distribution) then about 95% of the observations will
lie between the mean minus 2s and the mean plus 2s. From the histogram we observe that 600 to
1400 yards/min should include about 95% of the frequency distribution, so (1400 − 600)/4 = 200
yards/min. This is a very rough calculation!
15. (a) The mean should be k times larger.
(b) The standard deviation should be k times larger.
(c) The median should be k times larger.
(d) The interquartile range should be k times larger.
(e) The coefficient of variation will not change.
(f) The variance will be k2 times larger.
16. (a) There are not many observations, so it is difficult to say what the full distribution would look
like. Nevertheless, the point on the far left suggests that the distribution is strongly left-skewed
or perhaps has an outlier. The mean will be sensitive to the extreme observation, whereas the
median will not be affected. In this case the median is a better description of where the majority
of the data are located.
(b) The standard deviation is sensitive to extreme observations, whereas the interquartile is less
affected. In this case the interquartile range gives a better description of the spread of the bulk of
the data.
,17. (a) The frequency distributions are all right-skewed: The whiskers and span from median to third
quartile are greater than those on the opposite side of the box, and there are multiple extreme
values of actual survival times.
(b) The distributions for predictions of 6−24 months are broader (higher spread) than those for
predictions of 1−4 months, as indicated by a larger interquartile range.
(c) Median actual survival times increased slightly with increasing predicted survival times
between 1 to 6 months, but did not increase further for longer predicted survival times. Predicted
survival times tend to be over-optimistic: beyond predictions of about 2 months, median actual
survival times are consistently less than predicted times.
(d) The means will be greater than the medians because the distributions are right-skewed, and
so might be closer to the predicted survival times.
18. (a) Females had slightly higher mean LRS (1.7 recruits) than males (1.5 recruits).
(b) Every recruit must have both a father and a mother, so it is not easy to see why male and
female LRS should differ. One possibility is that females live longer than males. Another
possibility is that some females in the study mated with other males that were not part of the
sample.
(c) Females had slightly higher variance in LRS (4.3 recruits2) than males (3.5 recruits2).
Chapter 4
8. (a) SE = 6. = 0.10. In women, 4. = 0.06.
(b) Standard deviation, because it describes the spread of the distribution of the variable itself. In
contrast, the standard error describes the spread of the sampling distribution of the sample mean.
(c) The standard error, because it describes the spread of the distribution of sample means. If the
standard error is small, then the sample mean is likely close to the population mean (low
uncertainty).
(d) The study did not actually measure number of sexual partners, but merely reported the
number that respondents claimed. Perhaps men exaggerate their numbers or women
underestimate theirs. Another possibility is that men obtain partners also from women not
included in the survey (e.g., prostitutes or women living outside Britain).
9. (a) False.
(b) True.
(c) True.
(d) True.
10. No (the true mean and the sample confidence limits are all constants, so there is no probability
involved). The correct interpretation is that in 95% of random samples, the 95% confidence
interval calculated will contain the population mean.
11. (a) A histogram or cumulative frequency distribution.
(b) 8.3 genes.
(c) 0.7 genes.
(d) The spread of the sampling distribution of the mean number of genes regulated.
(e) That we have a random sample of the total population of regulatory genes.
12. (a) Using the 2SE method, 6.9 < < 9.8 genes.
, (b) The interval between 6.9 and 9.8 represents the most plausible values for the population
mean. In roughly 95% of random samples from the population, when we compute the 95%
confidence interval the interval will include the true population mean.
13. (a) False
(b) True.
(c) False.
(d) False.
Chapter 5
17. (a) No, some plants are tall with green pods, so "tall" and "green pods" are not mutually
exclusive.
(b) 1200/1600 were tall, 1200/1600 were green. If independent, Pr[tall and green] = Pr[tall]
Pr[green] = 3/4 3/4 = 9/16, or 900 out of 1600. There were 900 out of 1600, so it appears that
green and tall are independent.
18. (a) There are four kings, so Pr[draw king] = 4/52 = 1/13
(b) Pr[spade face card] = 3/52 (J, Q, K of spades = 3; 52 total)
(c) Pr[card without number] = A, K, Q, J of any suit = 16 /52.
(d) Pr[red] = 0.5 (13 diamonds, 13 hearts, 26 total). Pr[ace] = 4/52 = 1/13. Pr[red ace] = 2/52 =
1/26. There are red aces, so these are not mutually exclusive. Pr[red ace] = Pr[red] Pr[ace], so
they are independent.
(e) Mutually exclusive events include red or black; Jack or number; spade or diamond (and many
others).
(f) Pr[red king] = 2/52 = 1/26. Pr[face card hearts] = 3/52. No, these events are not mutually
exclusive: the king of hearts is a red king and a face card in hearts. No, these events are not
independent. This is easily shown by example: Pr[king hearts] = 1/52; this is not equal to Pr[red
king] Pr[face card hearts] = .
19. (a) If you pick any nucleotide from the first region, there is a 25% probability that you will pick
the same nucleotide from the second region, as all nucleotides have the same probability there.
Therefore, the odds that a random draw of one nucleotide from each region will match is 0.25.
(b) For a codon from the first region to match a codon from the second region, this is equivalent
to three independent draws occurring, each one matched in the two regions. If the probability of
one matching is 0.25, then the probability of three in a row matching is 1/43 = 0.015625.
20. (a) Pr[rain on random day in Vancouver] = (0.25 0.58) + (0.25 0.38) + (0.25 0.25) + (0.25
0.53) = 0.435.
(b) Pr[winter| raining] = Pr[raining| winter] Pr[winter] / Pr[raining] = 0.58 0..435 =
0.333
, 21. (a)
(b) Pr["Yes"] = (0.5 0.5) + (0.5 0.2) = 0.35
22. Pr[10 adenines in a row], if nucleotides are random in sequence and only draw 10, = 0.2510 =
9.54 10-7.
23. Imagine that the order of expression of eight genes is 1 to 8. What is the probability that these
eight genes would end up in this order on a chromosome if distributed randomly? There is a 1/8
chance that gene 1 will be first. If this is true, there is a 1/7 chance that gene 2 is second. If both
of these are true, there is a 1/6 chance that gene 3 is third, and so on. Overall, the product of the
independent probabilities is 2.48 10-5. This is an acceptable answer. However, the sequence of
genes would be in the same order as their expression starting at either the first or the eight gene,
so we should multiply this by two to get 4.96 10-5.
24. Pr[all land on chromosome A] = 0.18= 1 10-8. There are ten different chromosomes that they
could all land on, so 10 different ways to end up on the same chromosome. 10 10-8=10-7.
25. (a)
Overall probability of survival = (0.3 0.8) + (0.2 0.3) + (0.5 0.1) = 0.35
(b) Pr[Survival | lands] = 0.35.
(c) Pr[Survival] = Pr[Survival | lands] Pr[lands] = 0.35 0.8 = 0.28
26. (a) Pr[drawn pebble is white] = 2/5
(b) Pr[drawn pebble is white | first drawn is black] = 1/2
(c) Pr[three draws with replacement are white] = 0.43 = 0.064
(d) Pr[three sequential draws without replacement white] = 0 (there are only two white pebbles
in the bag!]