Instructor Solution Manual for Statistics for Business and Economics, 14th edition
richard@qwconsultancy.com
1|Pa ge
INSTRUCTOR’S SOLUTIONS MANUAL MARK DUMMELDINGER University of South Florida
S TATISTICS FOR B USINESS AND E CONOMICS
FOURTEENTH EDITION
James T. McClave University of Florida
P. George Benson College of Charleston
Terry Sincich University of South Florida
Please contact https://support.pearson.com/getsupport/s/contactsupport with any queries on this content. Microsoft and/or its respective suppliers make no representations about the suitability of the information contained in the documents and related graphics published as part of the services for any purpose. All such documents and related graphics are provided “as is” without warranty of any kind. Microsoft and/or its respective suppliers hereby disclaim all warranties and conditions with regard to this information, including all warranties and conditions of merchantability, whether express, implied or statutory, fitness for a particular purpose, title and non-infringement. In no event shall Microsoft and/or its respective suppliers be liable for any special, indirect or consequential damages or any damages whatsoever resulting from loss of use, data or profits, whether in an action of contract, negligence or other tortious action, arising out of or in connection with the use or performance of information available from the services. The documents and related graphics contained herein could include technical inaccuracies or typographical errors. Changes are periodically added to the information herein. Microsoft and/or its respective suppliers may make improvements and/or changes in the product(s) and/or the program(s) described herein at any time. Partial screen shots may be viewed in full within the software version specified. Microsoft® and Windows® are registered trademarks of the Microsoft Corporation in the U.S.A. and other countries. This book is not sponsored or endorsed by or affiliated with the Microsoft Corporation. The author and publisher of this book have used their best efforts in preparing this book. These efforts include the development, research, and testing of the theories and programs to determine their effectiveness. The author and publisher make no warranty of any kind, expressed or implied, with regard to these programs or the documentation contained in this book. The author and publisher shall not be liable in any event for incidental or consequential damages in connection with, or arising out of, the furnishing, performance, or use of these programs. Reproduced by Pearson from electronic files supplied by the author. Copyright © 2022, 2018, 2014 by Pearson Education, Inc. or its affiliates, 221 River Street, Hoboken, NJ 07030. All Rights Reserved. Manufactured in the United States of America. This publication is protected by copyright, and permission should be obtained from the publisher prior to any prohibited reproduction, storage in a retrieval system, or transmission in any form or by any means, electronic, mechanical, photocopying, recording, or otherwise. For information regarding permissions, request forms, and the appropriate contacts within the Pearson Education Global Rights and Permissions department, please visit www.pearsoned.com/permissions/. PEARSON, ALWAYS LEARNING, and MYLAB are exclusive trademarks owned by Pearson Education, Inc. or its affiliates in the U.S. and/or other countries. Unless otherwise indicated herein, any third-party trademarks, logos, or icons that may appear in this work are the property of their respective owners, and any references to third-party trademarks, logos, icons, or other trade dress are for demonstrative or descriptive purposes only. Such references are not intended to imply any sponsorship, endorsement, authorization, or promotion of Pearson’s products by the owners of such marks, or any relationship between the owner and Pearson Education, Inc., or its affiliates, authors, licensees, or distributors.
ISBN-13: 978-0-13-685527-9 ISBN-10: 0-13-685527-X
Contents 1. Statistics, Data, and Statistical Thinking 1 2. Methods for Describing Sets of Data 10 3. Probability 90 4. Random Variables and Probability Distributions 138 5. Sampling Distributions 235 6. Inferences Based on a Single Sample: Estimation with Confidence Intervals 268 7. Inferences Based on a Single Sample: Tests of Hypotheses 311 8. Inferences Based on Two Samples: Confidence Intervals and Tests of Hypotheses 368 9. Design of Experiments and Analysis of Variance 431 10. Categorical Data Analysis 506 11. Simple Linear Regression 546 12. Multiple Regression and Model Building 623 13. Methods for Quality Improvement: Statistical Process Control 745 14. Time Series: Descriptive Analyses, Models, and Forecasting 810 15. Nonparametric Statistics 883
Chapter 1 Statistics, Data, and Statistical Thinking 1.1
Statistics is a science that deals with the collection, classification, analysis, and interpretation of information or data. It is a meaningful, useful science with a broad, almost limitless scope of applications to business, government, and the physical and social sciences.
1.2
Descriptive statistics utilizes numerical and graphical methods to look for patterns, to summarize, and to present the information in a set of data. Inferential statistics utilizes sample data to make estimates, decisions, predictions, or other generalizations about a larger set of data.
1.3
The four elements of a descriptive statistics problem are: 1. 2. 3. 4.
1.4
The population or sample of interest. This is the collection of all the units upon which the variable is measured. One or more variables that are to be investigated. These are the types of data that are to be collected. Tables, graphs, or numerical summary tools. These are tools used to display the characteristic of the sample or population. Identification of patterns in the data. These are conclusions drawn from what the summary tools revealed about the population or sample.
The five elements of an inferential statistical analysis are: 1. 2. 3. 4. 5.
The population of interest. The population is a set of existing units. One or more variables that are to be investigated. A variable is a characteristic or property of an individual population unit. The sample of population units. A sample is a subset of the units of a population. The inference about the population based on information contained in the sample. A statistical inference is an estimate, prediction, or generalization about a population based on information contained in a sample. A measure of reliability for the inference. The reliability of an inference is how confident one is that the inference is correct.
1.5
The first major method of collecting data is from a published source. These data have already been collected by someone else and are available in a published source. The second method of collecting data is from a designed experiment. These data are collected by a researcher who exerts strict control over the experimental units in a study. These data are measured directly from the experimental units. The final method of collecting data is observational. These data are collected directly from experimental units by simply observing the experimental units in their natural environment and recording the values of the desired characteristics. The most common type of observational study is a survey.
1.6
Quantitative data are measurements that are recorded on a meaningful numerical scale. Qualitative data are measurements that are not numerical in nature; they can only be classified into one of a group of categories.
1.7
A population is a set of existing units such as people, objects, transactions, or events. A variable is a characteristic or property of an individual population unit such as height of a person, time of a reflex, amount of a transaction, etc. 1 Copyright © 2022 Pearson Education, Inc.
8
Chapter 1 2.
Using a scale from 1 to 5, where 1 means strongly disagree and 5 means strongly agree, indicate your agreement to the following statement: "The trend of consolidation in the banking industry will continue in the next five years." 1 strongly disagree
1.36
2 disagree
3 no opinion
4 agree
5 strongly agree
b.
The population of interest is the set of all bank presidents in the United States.
c.
It would be extremely difficult and costly to obtain information from all bank presidents. Thus, it would be more efficient to sample just 200 bank presidents. However, by sending the questionnaires to only 200 bank presidents, one risks getting the results from a sample which is not representative of the population. The sample must be chosen in such a way that the results will be representative of the entire population of bank presidents in order to be of any use.
Answers will vary. Using MINITAB, the 5 seven-digit phone numbers generated with area code 373 were: 373-639-0598 373-411-9164 373-502-7699 373-782-2719 373-930-3231
1.37
1.38
1.39
a.
The population of interest is the set of all people in the United States at least 15 years of age.
b.
The variable being measured is the employment status of each person. This variable is qualitative. Each person is either employed or not.
c.
The problem of interest to the Census Bureau is inferential. Based on the information contained in the sample, the Census Bureau wants to estimate the percentage of all people in the labor force who are unemployed.
a.
The process being studied is the process of filling beverage cans with softdrink at CCSB's Wakefield plant.
b.
The variable of interest is the amount of carbon dioxide added to each can of beverage.
c.
The sampling plan was to monitor five filled cans every 15 minutes. The sample is the total number of cans selected.
d.
The company's immediate interest is learning about the process of filling beverage cans with softdrink at CCSB's Wakefield plant. To do this, they are measuring the amount of carbon dioxide added to a can of beverage to make an inference about the process of filling beverage cans. In particular, they might use the mean amount of carbon dioxide added to the sampled cans of beverage to estimate the mean amount of carbon dioxide added to all the cans on the process line.
e.
The technician would then be dealing with a population. The cans of beverage have already been processed. He/she is now interested in the outputs.
Suppose we want to select 900 intersections by numbering the intersections from 1 to 500,000. We would then use a random number table or a random number generator from a software program to select 900 distinct intersection points. These would then be the sampled markets. Now, suppose we want to select the 900 intersections by selecting a row from the 500 and a column from the 1,000. We would first number the rows from 1 to 500 and number the columns from 1 to 1,000. Using Copyright © 2022 Pearson Education, Inc.
Statistics, Data, and Statistical Thinking
9
a random number generator, we would generate a sample of 900 from the 500 rows. Obviously, many rows will be selected more than once. At the same time, we use a random number generator to select 900 columns from the 1,000 columns. Again, some of the columns could be selected more than once. Placing these two sets of random numbers side-by-side, we would use the row-column combinations to select the intersections. For example, suppose the first row selected was 453 and the first column selected was 731. The first intersection selected would be row 453, column 731. This process would be continued until 900 unique intersections were selected. 1.40
Answers will vary. a.
The results as stated indicate that by eating oat bran, one can improve his/her health. However, the only way to get the stated benefit is to eat only oat bran with limited results. People may change their eating habits expecting an outcome that is almost impossible.
b.
In order to investigate the impact of domestic violence on birth defects, one would need to collect data on all kinds of birth defects and whether the mother suffered any domestic violence or not during her pregnancy. One could use an observational study survey to collect the data.
c.
Very few people are always happy with the way they are. However, many people are happy with themselves most of the time. One might want to ask a series of questions to measure self-esteem rather than just one. One question might ask what percent of the time the high school girl is happy with the way she is.
d.
The results of the study are probably misleading because of the fact that if someone relied on a limited number of foods to feed her children it does not imply that the children are hungry. In addition, one might cut the size of a meal because the children were overweight, not because there was not enough food. One might get better information about the proportion of hungry American children by actually recording what a large, representative sample of children eat in a week.
e.
A leading question gives information that seems to be true, but may not be complete. Based on the incomplete information, the respondent may come to a different decision than if the information was not provided.
Copyright © 2022 Pearson Education, Inc.
Chapter 2 Methods for Describing Sets of Data 2.1
First, we find the frequency of the grade A. The sum of the frequencies for all five grades must be 200. Therefore, subtract the sum of the frequencies of the other four grades from 200. The frequency for grade A is: 200 − (36 + 90 + 30 + 28) = 200 − 184 = 16 To find the relative frequency for each grade, divide the frequency by the total sample size, 200. The relative frequency for the grade B is 36/200 = .18. The rest of the relative frequencies are found in a similar manner and appear in the table: Grade on Statistics Exam A: 90 −100 B: 80 − 89 C: 65 − 79 D: 50 − 64 F: Below 50 Total
2.2
a.
Relative Frequency .08 .18 .45 .15 .14 1.00
To find the frequency for each class, count the number of times each letter occurs. The frequencies for the three classes are: Class X Y Z Total
b.
Frequency 16 36 90 30 28 200
Frequency 8 9 3 20
The relative frequency for each class is found by dividing the frequency by the total sample size. The relative frequency for the class X is 8/20 = .40. The relative frequency for the class Y is 9/20 = .45. The relative frequency for the class Z is 3/20 = .15. Class X Y Z Total
Frequency 8 9 3 20
Relative Frequency .40 .45 .15 1.00
10 Copyright © 2022 Pearson Education, Inc.
Methods for Describing Sets of Data c.
The frequency bar chart is:
9 8
Frequency
7 6 5 4 3 2 1 0
d.
X
Y C la s s
Z
The pie chart for the frequency distribution is: Pie Chart of Class Category X Y Z
Z 15.0%
X 40.0%
Y 45.0%
2.3
a.
The bar graph for the male student data is shown here:
Copyright © 2022 Pearson Education, Inc.
11
58
Chapter 2 ̄
𝑧=
⇒ −3 =
.
⇒ −3(3.9725) = 𝑥 − 94.882 ⇒ −11.918 = 𝑥 − 94.882 ⇒ 𝑥 = 82.965
.
Observations greater than 1046.800 or less than 82.965 would be considered outliers. Using this criterion, the following observations would be outliers: 79 and 79.
2.118
c.
Yes, these methods do agree exactly. Both methods identify two observations as outliers.
a.
Using MINITAB, the box plot is: Boxplot of Downtime
0
10
20
30
40
50
60
70
Downtime
The median is about 18. The data appear to be skewed to the right since there are 3 suspect outliers to the right and none to the left. The variability of the data is fairly small because the IQR is fairly small, approximately 26 − 10 = 16. b.
The customers associated with the suspected outliers are customers 268, 269, and 264.
c.
In order to find the z-scores, we must first find the mean and standard deviation. 𝑥̄ =
∑
=
s2 =
= 20.375
( x) x − 2
n
n −1
2
𝑠 = √192.90705 = 13.89 The z-scores associated with the suspected outliers are: Customer 268 𝑧 = Customer 269 𝑧 = Customer 264 𝑧 =
.
= 2.06
.
= 2.13
.
= 3.14
.
.
.
2
24129 − 815 40 = 192.90705 = 40 − 1
All the z-scores are greater than 2. These are unusual values.
Copyright © 2022 Pearson Education, Inc.
Methods for Describing Sets of Data 2.119
59
Using MINITAB, the boxplots of the data are: Boxplot of PermA, PermB, PermC
PermA
PermB
PermC
50
75
100
125
150
Data
The descriptive statistics are: Descriptive Statistics: PermA, PermB, PermC Variable PermA PermB PermC
2.120
N 100 100 100
Mean 73.62 128.54 83.07
StDev 14.48 21.97 20.05
Minimum 55.20 50.40 52.20
Q1 62.00 108.65 67.72
Median 70.45 139.30 78.65
Q3 81.42 147.02 95.35
Maximum 122.40 150.00 129.00
IQR 19.42 38.37 27.63
a.
For group A, the suspect outliers are any observations greater than𝑄 + 1.5(𝐼𝑄𝑅) = 81.42 + 1.5(19.42) = 110.55or less than𝑄 − 1.5(𝐼𝑄𝑅) = 62 − 1.5(19.42) = 32.87. There are 3 observations greater than 110.55: 117.3, 118.5, and 122.4.
b.
For group B, the suspect outliers are any observations greater than𝑄 + 1.5(𝐼𝑄𝑅) = 147.02 + 1.5(38.37) = 204.575or less than𝑄 − 1.5(𝐼𝑄𝑅) = 108.65 − 1.5(38.37) = 51.095. There is 1 observation less than 51.095: 50.4.
c.
For group C, the suspect outliers are any observations greater than𝑄 + 1.5(𝐼𝑄𝑅) = 95.35 + 1.5(27.63) = 136.795or less than𝑄 − 1.5(𝐼𝑄𝑅) = 67.72 − 1.5(27.63) = 26.275. No observations are greater than 136.795 or less than 26.275.
d.
For group A, if the outliers are removed, the mean will decrease, the median will slightly decrease, and the standard deviation will decrease. For group B, if the outlier is removed, the mean will increase, the median will slightly increase, and the standard deviation will decrease.
For Perturbed Intrinsics, but no Perturbed Projections: 𝑥̄ =
∑
=
.
= 1.62
𝑠 =
∑
∑
=
.
.
̄
The z-score corresponding to a value of 4.5 is𝑧 =
.
= =
.
.
= .627
𝑠 = √𝑠 = √. 627 = .792
= 3.63
.
Since this z-score is greater than 3, we would consider this an outlier for perturbed intrinsics, but no perturbed projections. For Perturbed Projections, but no Perturbed Intrinsics: 𝑥̄ =
∑
=
.
= 25.16 𝑠 =
∑
∑
=
.
.
=
𝑠 = √𝑠 = √46.243 = 6.800 Copyright © 2022 Pearson Education, Inc.
.
= 46.243
60
Chapter 2
The z-score corresponding to a value of 4.5 is𝑧 =
̄
=
.
.
= −3.038
.
Since this z-score is less than -3, we would consider this an outlier for perturbed projections, but no perturbed intrinsics. Since the z-score corresponding to 4.5 for the perturbed projections, but no perturbed intrinsics is smaller in absolute value than that for perturbed intrinsics, but no perturbed projections, it is more likely that the that the type of camera perturbation is perturbed projections, but no perturbed intrinsics. 2.121
From the stem-and-leaf display in Exercise 2.34, the data are fairly mound-shaped, but skewed somewhat to the right. The sample mean is𝑥̄ =
∑
=
The sample variance is𝑠 =
∑
= 59.72. (∑ )
=
,
= 321.7933.
The sample standard deviation is𝑠 = √321.7933 = 17.9386. The z-score associated with the largest value is𝑧 =
̄
.
=
.
= 2.36.
This observation is a suspect outlier. The observations associated with the one-time customers are 5 of the largest 7 observations. Thus, repeat customers tend to have shorter delivery times than one-time customers. Using MINITAB, a scatterplot of the data is: Scatterplot of Var 2 vs Var 1 14 12 10
Var 2
2.122
8 6 4 2 0 0
2
4
6
8
Var 1
Copyright © 2022 Pearson Education, Inc.
Methods for Describing Sets of Data 2.123
61
Using MINITAB, the scatterplot is: Scatterplot of Var 2 vs Var 1 18 16 14
Var 2
12 10 8 6 4 2 0 1
2
3
4
5
Var 1
2.124
Using MINITAB, the scatterplot is: Scatterplot of RATIO vs DIAMETER 10.0 9.5
RATIO
9.0 8.5 8.0 7.5 7.0 6.5 0
100
200
300
400
500
600
700
DIAMETER
It appears that as the pipe diameter increases, the ratio of repair to replacement cost increases. 2.125.
From the scatterplot of the data, it appears that as the number of punishments increases, the average payoff decreases. Thus, there appears to be a negative linear relationship between punishment use and average payoff. This supports the researchers conclusion that “winners” don’t punish”.
Copyright © 2022 Pearson Education, Inc.
106
Chapter 3 9 126 135 144 225 234 333
# ways 6 6 3 3 6 _ 1 25
10 136 145 226 235 244 334
# ways 6 6 3 6 3 _3 27
Thus, there are a total of 25 ways to get a sum of 9 and 27 ways to get a sum of 10. The chance of throwing a sum of 9 (25 chances out of 216 possibilities) is less than the chance of throwing a 10 (27 chances out of 216 possibilities). 3.51
A possible Venn Diagram would be:
ACC Dimension
CCC
Plain
Lambda
3.52
3.53
3.54
CFA
a.
P( A ∩ B ) = P ( A | B) P ( B) = .6(.2) = .12
b.
P( B | A) =
P( A ∩ B) .12 = = .3 P( A) .4
a.
P( A | B) =
P( A ∩ B) .1 = = .5 .2 P( B)
b.
P( B | A) =
P( A ∩ B) .1 = = .25 P( A) .4
c.
Events A and B are said to be independent if P( A | B) = P( A) . In this case, P ( A | B ) = .5 and P( A) = .4 . Thus, A and B are not independent.
a.
Since A and B are mutually exclusive events, P( A ∪ B) = P( A) + P( B) = .30 + .55 = .85
b.
Since A and C are mutually exclusive events, P( A ∩ C ) = 0
c.
P( A | B) =
P( A ∩ B) 0 = =0 P( B) .55 Copyright © 2022 Pearson Education, Inc.
Probability
3.55
3.56
d.
Since B and C are mutually exclusive events, P( B ∪ C ) = P( B) + P(C ) = .55 + .15 = .70
e.
No, B and C cannot be independent events because they are mutually exclusive events.
a.
If two events are independent, then P( A ∩ B) = P( A) P( B) = .4(.2) = .08 .
b.
If two events are independent, then P( A | B) = P( A) = .4 .
c.
P( A ∪ B) = P( A) + P( B) − P( A ∩ B) = .4 + .2 − .08 = .52
a.
If two fair coins are tossed, there are 4 possible outcomes or simple events. They are: E1 = HH
107
E2 = HT E3 = TH E4 = TT
Event A contains the simple events E1, E2, and E3. Event B contains the simple events E2 and E3. A Venn diagram of this would be:
A
B E2 E3
E1
E4
Since the coins are fair, each of the sample points is equally likely. Each would have probabilities of ¼. b.
1 3 P( A) = 3 = = .75 4 4 P ( A ∩ B ) = P ( E2 )+P ( E3 ) =
3.57
1 2 1 P( B) = 2 = = = .5 4 4 2 1 1 2 1 + = = = .5 4 4 4 2
P( A ∩ B) .5 = =1 P( B) .5
c.
P ( A | B) =
P( B | A) =
a.
P( A) = P( E1 ) + P ( E2 ) + P( E3 ) = .2 + .3 + .3 = .8
P( A ∩ B) .5 = = .667 P( A) .75
P( B) = P( E2 ) + P( E3 ) + P ( E5 ) = .3 + .3 + .1 = .7
Copyright © 2022 Pearson Education, Inc.
108
Chapter 3
P( A ∩ B) = P( E2 ) + P( E3 ) = .3 + .3 = .6 b.
P( E1 | A) =
P( E 1 ∩ A) P( E 1) .2 = = = .25 P( A) P( A) .8
P( E2 | A) =
P( E 2 ∩ A) P( E 2) .3 = = = .375 P ( A) P( A) .8
P( E3 | A) =
P( E 3 ∩ A) P( E 3) .3 = = = .375 P( A) P( A) .8
The original sample point probabilities are in the proportion .2 to .3 to .3 or 2 to 3 to 3. The conditional probabilities for these sample points are in the proportion .25 to .375 to .375 or 2 to 3 to 3. c.
(1)
P( B | A) = P( E2 | A) + P( E3 | A) = .375 + .375 = .75 (from part b)
(2)
P( B | A) =
P( A ∩ B) .6 = = .75 (from part a) P( A) .8
The two methods do yield the same result. d.
3.58
If A and B are independent events, P( B | A) = P( B) . From part c, P ( B | A) = .75 . From part a, P( B) = .7 . Since .75 ≠ .7 , A and B are not independent events.
The 36 possible outcomes obtained when tossing two dice are listed below: (1, 1) (1, 2) (1, 3) (1, 4) (1, 5) (1, 6) (2, 1) (2, 2) (2, 3) (2, 4) (2, 5) (2, 6) (3, 1) (3, 2) (3, 3) (3, 4) (3, 5) (3, 6) (4, 1) (4, 2) (4, 3) (4, 4) (4, 5) (4, 6) (5, 1) (5, 2) (5, 3) (5, 4) (5, 5) (5, 6) (6, 1) (6, 2) (6, 3) (6, 4) (6, 5) (6, 6) A: {(1, 2), (1, 4), (1, 6), (2, 1), (2, 3), (2, 5), (3, 2), (3, 4), (3, 6), (4, 1), (4, 3), (4, 5), (5, 2), (5, 4), (5, 6), (6, 1), (6, 3), (6, 5)} B: {(3, 6), (4, 5), (5, 4), (5, 6), (6, 3), (6, 5), (6, 6)} A ∩ B : {(3, 6), (4, 5), (5, 4), (5, 6), (6, 3), (6, 5)}
If A and B are independent, then P( A) P( B) = P( A ∩ B) . P ( A) =
18 1 = 36 2
P ( A) P ( B ) =
P( B) =
7 36
P( A ∩ B) =
6 1 = 36 6
1 7 7 1 ⋅ = ≠ = P ( A ∩ B ) . Thus, A and B are not independent. 2 36 72 6
Copyright © 2022 Pearson Education, Inc.
Probability
3.59
a.
P ( A ∩ C ) = 0 A and C are mutually exclusive. P ( B ∩ C ) = 0 B and C are mutually exclusive.
b.
P( A) = P(1) + P(2) + P(3) = .20 + .05 + .30 = .55
P( B) = P(3) + P(4) = .30 + .10 = .40
P(C ) = P(5) + P(6) = .10 + .25 = .35
P( A ∩ B) = P(3) = .30
P( A | B) =
P( A ∩ B) .30 = = .75 P( B) .40
A and B are independent if P( A | B) = P( A) . Since P ( A | B ) = .75 and P ( A) = .55 , A and B are not independent. Since A and C are mutually exclusive, they are not independent. Similarly, since B and C are mutually exclusive, they are not independent. c.
Using the probabilities of sample points, P( A ∪ B) = P(1) + P(2) + P(3) + P(4) = .20 + .05 + .30 + .10 = .65
Using the additive rule, P( A ∪ B) = P( A) + P( B) − P( A ∩ B) = .55 + .40 − .30 = .65 Using the probabilities of sample points, P( A ∪ C ) = P(1) + P(2) + P(3) + P(5) + P(6) = .20 + .05 + .30 + .10 + .25 = .90 Using the additive rule, P( A ∪ C ) = P( A) + P(C ) − P( A ∩ C ) = .55 + .35 − 0 = .90 3.60
3.61
From the Exercise, P( A) = .15 , P( B) = .10 , and P ( A ∩ B) = .05 . a.
If events A and B are mutually exclusive then P( A ∩ B) = 0 . For this problem, P ( A ∩ B ) = .05 . Therefore, events A and B are not mutually exclusive.
b.
P( B | A) =
c.
Events A and B are independent if P( B | A) = P( B) . For this exercise, P ( B | A) = .333 and P( B) = .10 . Since these are not equal, events A and B are not independent.
P( A ∩ B) .05 = = .333 P( A) .15
Define the following events: A: {American believes the American Dream is within reach} F: {Person is female} From the problem, we know that 𝑃(𝐴) = 0.24 and 𝑃(𝐹|𝐴) = 0.63 𝑃(𝐴 ∩ 𝐹) = 𝑃(𝐴)𝑃(𝐹|𝐴) = 0.24(0.63) = 0.1512.
Copyright © 2022 Pearson Education, Inc.
109
196 4.129
Chapter 4 We will look at the 4 methods or determining if the 3 variables are normal. Distance: First, we will look at A histogram of the data. Using MINITAB, the histogram of the distance data is:
From the histogram, the distance data does appear to have a normal distribution. Next, we look at the intervals 𝑥̄ ± 𝑠, 𝑥̄ ± 2𝑠, 𝑥̄ ± 3𝑠. If the proportions of observations falling in each interval are approximately .68, .95, and 1.00, then the data are approximately normal. Using MINITAB, the summary statistics are: Statistics Variable
Total Count
DISTANCE ACCURACY INDEX
Mean StDev Minimum
25 305.73 10.63 25 61.36 7.20 25 3.171 1.984
Q1 Median
282.00 299.85 48.79 56.70 0.565 1.722
Q3 Maximum
305.60 314.30 60.32 65.94 2.917 3.929
327.10 80.95 9.434
IQR
14.45 9.25 2.207
For Distance: 𝑥̄ ± 𝑠 ⇒ 305.73 ± 10.63 ⇒ (295.10, 316.36). 16 of the 25 values fall in this interval. The proportion is 16/25 = .64. This is fairly close to the .68 we would expect if the data were normal. 𝑥̄ ± 2𝑠 ⇒ 305.73 ± 2(10.63) ⇒ 305.73 ± 21.64 ⇒ (284.47, 326.99). 23 of the 25 values fall in this interval. The proportion is 23/25 = .92. This is close to the .95 we would expect if the data were normal. 𝑥̄ ± 3𝑠 ⇒ 305.73 ± 3(10.63) ⇒ 305.73 ± 31.89 ⇒ (273.84, 337.62). 25 of the 25 values fall in this interval. The proportion is 25/25 = 1.00. This is equal to the 1.00 we would expect if the data were normal. From this method, it appears that the distance data may be normal. Next, we look at the ratio of the IQR to s. .
= = 1.36 This is very close to the 1.3 we would expect if the data were normal. This method . indicates the distance data may be normal.
Copyright © 2022 Pearson Education, Inc.
Random Variables and Probability Distributions
197
Finally, using MINITAB, the normal probability plot is:
Since the data do form a fairly straight line, the distance data may be normal. From the 4 different methods, all indications are that the distance data are normal. Accuracy: First, we will look at a histogram of the data. Using MINITAB, the histogram of the accuracy data is:
From the histogram, the accuracy data do appear to have a fairly normal distribution. From the Descriptive Statistics above:
𝑥̄ ± 𝑠 ⇒ 61.36 ± 7.20 ⇒ (54.16, 68.56) 19 of the 25 values fall in this interval. The proportion is 19/25 = .76. This is greater than the .68 we would expect if the data were normal. 𝑥̄ ± 2𝑠 ⇒ 61.36 ± 2(7.20) ⇒ 61.36 ± 14.40 ⇒ (46.96, 75.76) 24 of the 25 values fall in this interval. The proportion is 24/25 = .96 This is very close to the .95 we would expect if the data were normal. 𝑥̄ ± 3𝑠 ⇒ 61.36 ± 3(7.20) ⇒ 61.36 ± 21.60 ⇒ (39.76, 82.96) 25 of the 25 values fall in this interval. The proportion is 25/25 = 1.00. This is equal to the 1.00 we would expect if the data were normal. From this method, it appears that the accuracy data may be normal. Next, we look at the ratio of the IQR to s.
Copyright © 2022 Pearson Education, Inc.
198
Chapter 4 .
= = 1.28. This is fairly close to the 1.3 we would expect if the data were normal. This method . indicates the accuracy data may be normal. Finally, using MINTAB, the normal probability plot is:
Since the data do form a fairly straight line, the accuracy data may be normal. From the 4 different methods, all indications are that the accuracy data might be normal. Index: First, we will look at a histogram of the data. Using MINITAB, the histogram of the index data is:
From the histogram, the index data do not appear to have a normal distribution. From the Descriptive Statistics above:
𝑥̄ ± 𝑠 ⇒ 3.171 ± 1.984 ⇒ (1.187, 5.155). 19 of the 25 values fall in this interval. The proportion is 19/25 = .76. This is greater than the .68 we would expect if the data were normal. 𝑥̄ ± 2𝑠 ⇒ 3.171 ± 2(1.984) ⇒ 3.171 ± 3.968 ⇒ (−0.797, 7.139). 24 of the 25 values fall in this interval. The proportion is 24/25 = .96. This is very close to the .95 we would expect if the data were normal.
Copyright © 2022 Pearson Education, Inc.
Random Variables and Probability Distributions
199
𝑥̄ ± 3𝑠 ⇒ 3.171 ± 3(1.984) ⇒ 3.171 ± 5.952 ⇒ (−2.781, 9.123). 25 of the 25 values fall in this interval. The proportion is 25/25 = 1.000. This is equal to the 1.00 we would expect if the data were normal. From this method, it appears that the index data may not be normal. Next, we look at the ratio of the IQR to s. .
= = 1.11. This is not very close to the 1.3 we would expect if the data were normal. This method . indicates the index data may not be normal. Finally, using MINTAB, the normal probability plot is:
Since the data do not form a fairly straight line, the index data may not be normal. From 3 of the 4 different methods, the indications are that the index data are not normal. Using MINITAB, the histograms of the data are: Histogram of PermA, PermB, PermC Normal
PermA
20
PermB
PermA Mean 73.62 StDev 14.48 N 100 PermB Mean 128.5 StDev 21.97 N 100
40 15
30
10
Frequency
4.130
20
5
10
0
0
45
60
75
90
105
120
60
80
100
120
140
160
180
PermC
16 12 8 4 0
45
60
75
90
105
120
Copyright © 2022 Pearson Education, Inc.
PermC Mean 83.07 StDev 20.05 N 100
200
Chapter 4 None of the three histograms appear to be mound-shaped. The histograms for Groups A and C appear to be skewed to the right, while the histogram for Group B appears to be skewed to the left. Thus, it does not appear that any of the 3 distributions are normally distributed. Next, we look at the intervals 𝑥̄ ± 𝑠, 𝑥̄ ± 2𝑠, 𝑥̄ ± 3𝑠. If the proportions of observations falling in each interval are approximately .68, .95, and 1.00, then the data are approximately normal. Using MINITAB, the summary statistics are: Descriptive Statistics: PermA, PermB, PermC Variable PermA PermB PermC
N 100 100 100
Mean 73.62 128.54 83.07
StDev 14.48 21.97 20.05
Minimum 55.20 50.40 52.20
Q1 62.00 108.65 67.72
Median 70.45 139.30 78.65
Q3 81.42 147.02 95.35
Maximum 122.40 150.00 129.00
IQR 19.42 38.37 27.63
For Group A: 𝑥̄ ± 𝑠 ⇒ 73.62 ± 14.48 ⇒ (59.14, 88.10). 76 of the 100 values fall in this interval. The proportion is 76/100 = .76. This is much larger than the .68 we would expect if the data were normal. 𝑥̄ ± 𝑠 ⇒ 73.62 ± 2(14.48) ⇒ 73.62 ± 28.96 ⇒ (44.66, 102.58). 96 of the 100 values fall in this interval. The proportion is 96/100 = .96. This is slightly larger than the .95 we would expect if the data were normal. 𝑥̄ ± 𝑠 ⇒ 73.62 ± 3(14.48) ⇒ 73.62 ± 43.44 ⇒ (30.18, 117.06). 97 of the 100 values fall in this interval. The proportion is 97/100 = .97. This is much smaller than the 1.00 we would expect if the data were normal. From this method, it appears that the data are not normal. For Group B: 𝑥̄ ± 𝑠 ⇒ 128.54 ± 21.97 ⇒ (106.57, 150.51). 81 of the 100 values fall in this interval. The proportion is 81/100 = .81. This is much larger than the .68 we would expect if the data were normal. 𝑥̄ ± 𝑠 ⇒ 128.54 ± 2(21.97) ⇒ 128.54 ± 43.94 ⇒ (84.60, 172.48). 98 of the 100 values fall in this interval. The proportion is 98/100 = .98. This is larger than the .95 we would expect if the data were normal. 𝑥̄ ± 𝑠 ⇒ 128.54 ± 3(21.97) ⇒ 128.54 ± 65.91 ⇒ (62.63, 194.45). 98 of the 100 values fall in this interval. The proportion is 98/100 = .98. This is somewhat smaller than the 1.00 we would expect if the data were normal. From this method, it appears that the data are not normal. For Group C: 𝑥̄ ± 𝑠 ⇒ 83.07 ± 20.05 ⇒ (63.02, 103.12). 66 of the 100 values fall in this interval. The proportion is 66/100 = .66. This is about the same as the .68 we would expect if the data were normal. 𝑥̄ ± 𝑠 ⇒ 83.07 ± 2(20.05) ⇒ 83.07 ± 40.10 ⇒ (42.97, 123.17). 96 of the 100 values fall in this interval. The proportion is 96/100 = .96. This is slightly larger than the .95 we would expect if the data were normal. 𝑥̄ ± 𝑠 ⇒ 83.07 ± 3(20.05) ⇒ 83.07 ± 60.15 ⇒ (22.92, 143.22). 100 of the 100 values fall in this interval. The proportion is 100/100 = 1.00. This agrees with the 1.00 we would expect if the data were normal. From this method, it appears that the data are approximately normal.
Copyright © 2022 Pearson Education, Inc.
Random Variables and Probability Distributions
201
Next, we look at the ratio of the IQR to s. .
For Group A, = = 1.341. This is fairly close to the 1.3 we would expect if the data were normal. . This method indicates the data could be normal. .
For Group B, = = 1.746. This is larger than the 1.3 we would expect if the data were normal. . This method indicates the data may not be normal. .
For Group C, = = 1.378. This is fairly close to the 1.3 we would expect if the data were normal. . This method indicates the data could be normal. Finally, using MINITAB, the normal probability plot is: Probability Plot of PermA, PermB, PermC Normal - 95% CI
PermA
Percent
99.9 99 90
90
50
50
10
10
1
1
0.1
0.1
30
60
PermB
99.9 99
90
120
50
100
PermA Mean 73.62 StDev 14.48 N 100 AD 2.564 P-Value <0.005
150
PermB Mean 128.5 StDev 21.97 N 100 AD 5.887 P-Value <0.005
200
PermC
99.9 99
PermC Mean 83.07 StDev 20.05 N 100 AD 2.205 P-Value <0.005
90 50 10 1 0.1
0
40
80
120
160
The data do not form a straight line for any of the 3 groups. This indicates that the data are probably not normal. Thus, based on the histograms and normal probability plots, it appears that the data for all three groups are not normally distributed. 4.131
If the data are normally distributed, the distribution will be symmetric and the mean and median will be close in value. For this data set, the mean is much greater than the median, indicating the data are not normally distributed. Thus, it is very unlikely that the data are normally distributed.
4.132
a.
(𝑐 ≤ 𝑥 ≤ 𝑑)
𝑓(𝑥) = = 𝑓(𝑥) =
b.
𝜇=
= (3 ≤ 𝑥 ≤ 7) 0 otherwise =
=
=5
𝜎=
√
=
√
=
√
= 1.155
Copyright © 2022 Pearson Education, Inc.
202
Chapter 4 c.
𝜇 ± 𝜎 ⇒ 5 ± 1.155 ⇒ (3.845, 6.155) =
𝑃(𝜇 − 𝜎 ≤ 𝑥 ≤ 𝜇 + 𝜎) = 𝑃(3.845 ≤ 𝑥 ≤ 6.155) = 4.133
a.
(𝑐 ≤ 𝑥 ≤ 𝑑)
𝑓(𝑥) = =
=
= .04
.04 (20 ≤ 𝑥 ≤ 45) 0 otherwise
So, 𝑓(𝑥) = b.
𝜇=
c.
Using MINITAB, the graph is:
f(x)
.
=
=
= 32.5
𝜎=
√
=
√
= 7.22
1/25
0
20
45 x
𝜇 − 2𝜎 = 18.06
𝜇 = 32.5
𝜇 + 2𝜎 = 46.94
𝜇 ± 2𝜎 ⇒ 32.5 ± 2(7.22) ⇒ (18.06, 46.94) 𝑃(18. 06 < 𝑥 < 46.94) = 𝑃(20 < 𝑥 < 45) = (45 − 20)(. 04) = 1 4.134
From Exercise 4.133, 𝑓(𝑥) =
.04 (20 ≤ 𝑥 ≤ 45) 0 otherwise
a.
𝑃(20 ≤ 𝑥 ≤ 30) = (30 − 20)(. 04) = .4
b.
𝑃(20 < 𝑥 ≤ 30) = (30 − 20)(. 04) = .4
c.
𝑃(𝑥 ≥ 30) = (45 − 30)(. 04) = .6
d.
𝑃(𝑥 ≥ 45) = (45 − 45)(. 04) = 0
e.
𝑃(𝑥 ≤ 40) = (40 − 20)(. 04) = .8
f.
𝑃(𝑥 < 40) = (40 − 20)(. 04) = .8
g.
𝑃(15 ≤ 𝑥 ≤ 35) = (35 − 20)(.04) = .6
h.
𝑃(21. 5 ≤ 𝑥 ≤ 31. 5) = (31. 5 − 21. 5)(. 04) = .4 Copyright © 2022 Pearson Education, Inc.
.
=
.
= .5775
Random Variables and Probability Distributions 4.135
4.136
4.137
𝑃(𝑥 ≥ 𝑎) = 𝑒
/
/
=𝑒 /
. Using a calculator:
a.
𝑃(𝑥 > 1) = 𝑒
b.
𝑃(𝑥 ≤ 3) = 1 − 𝑃(𝑥 > 3) = 1 − 𝑒
c.
𝑃(𝑥 > 1.5) = 𝑒
d.
𝑃(𝑥 ≤ 5) = 1 − 𝑃(𝑥 > 5) = 1 − 𝑒
/
a.
𝑃(𝑥 ≤ 4) = 1 − 𝑃(𝑥 > 4) = 1 − 𝑒
/ .
b.
𝑃(𝑥 > 5) = 𝑒
c.
𝑃(𝑥 ≤ 2) = 1 − 𝑃(𝑥 > 2) = 1 − 𝑒
d.
𝑃(𝑥 > 3) = 𝑒
𝑓(𝑥) =
=𝑒
. /
/ .
/ .
=
= .367879
.
=𝑒
=𝑒
=1−𝑒
= 1 − .049787 = .950213
= .223130 =1−𝑒 =1−𝑒
= 1 − .006738 = .993262 .
= 1 − .201897 = .798103
= .135335
=𝑒
.
=
= .01
/ .
= 1 − 𝑒 . = 1 − .449329 = .550671
= .301194
𝑓(𝑥) =
.01 (100 ≤ 𝑥 ≤ 200) 0 𝑜𝑡ℎ𝑒𝑟𝑤𝑖𝑠𝑒
𝜇=
=
=
/
= 150
𝜎=
√
=
√
=
√
= 28.8675
a.
𝜇 ± 2𝜎 ⇒ 150 ± 2(28.8675) ⇒ 150 ± 57.735 ⇒ (92.265, 207.735) 𝑃(𝑥 < 92.265) + 𝑃(𝑥 > 207.735) = 𝑃(𝑥 < 100) + 𝑃(𝑥 > 200) = 0 + 0 = 0
b.
𝜇 ± 3𝜎 ⇒ 150 ± 3(28.8675) ⇒ 150 ± 86.6025 ⇒ (63.3975, 236.6025) 𝑃(63.3975 < 𝑥 < 236. 6025) = 𝑃(100 < 𝑥 < 200) = (200 − 100)(. 01) = 1
c.
From a, 𝜇 ± 2𝜎 ⇒ (92.265, 207.735). 𝑃(92.265 < 𝑥 < 207.735) = 𝑃(100 < 𝑥 < 200) = (200 − 100)(. 01) = 1
4.138
With 𝜃=2, 𝜇 = 𝜎 = 𝜃 = 2 a.
𝜇 ± 3𝜎 ⇒ 2 ± 3(2) ⇒ 2 ± 6 ⇒ (−4, 8) Since 𝜇 − 3𝜎 lies below 0, find the probability that x is more than 𝜇 + 3𝜎 = 8. 𝑃(𝑥 > 8) = 𝑒
b.
/
=𝑒
= .018316
𝜇 ± 2𝜎 ⇒ 2 ± 2(2) ⇒ 2 ± 4 ⇒ (−2, 6) Since 𝜇 − 2𝜎 lies below 0, find the probability that x is between 0 and 6. 𝑃(𝑥 < 6) = 1 − 𝑃(𝑥 ≥ 6) = 1 − 𝑒
/
= 1 − 𝑒 = 1 − .049787 = .950213 (using Table V, Appendix D)
Copyright © 2022 Pearson Education, Inc.
203