ELECTIVE IV - DIGITAL SKILLS & GRAPHIC REPRESENTATION (Course Code - 19MEL-6CR22S)
Faculty In-charge : Mr. Antim Dev Mishra
MID & END TERM SUBMISSION Submitted By : PARAS MONGIA Roll No. : 200MPLUP021 MASTERS IN URBAN PLANNING 2nd Year, 4th SEMESTER School Of Planning & Development (SPD)
CONTENT OF ASSIGNMENTS TASKS
TITLE
DATE
-
Introduction to Data Analytics and Visualization Types of Data, Data Statistics, Statistical Parameters
16.02.2022
SIGNATURE
1-2
To prepare a marksheet and then calculate:
1 2
3
Maxima, Minima, Average and median of the given data of students
Using Data Analysis tools to calculate: Mean, Median, Mode and Standard Deviation
3
02.03.2022
09.03.2022
4
16.03.2022
5-6
23.03.2022
7
To calculate the grade sheet of 10 students using MS Excel formulae or function (IF) with following conditions: = or > 80% - Grade A = or > 70% - Grade B = or > 60% - Grade C = or > 50% - Grade D
Analysing Data using :
PAGE NO.
4
Countif, Countifs, Sumif, Sumifs, Concatenate, Len
5
Calculate and Interpret Chi Square in SPSS
20.04.2022
8
6
Calculate and Interpret One Sample T Test in SPSS
27.04.2022
9
7
Calculate and Interpret Independent T Test in SPSS
11.05.2022
10
ASSIGNMENT 1 – MID TERM
ASSIGNMENT 2 – END TERM
ASSIGNMENT 1- BRIEF The Difference Between Data and Statistics
I N T R O D U C T I O N
While the terms ‘data’ and ‘statistics’ are often used interchangeably, in scholarly research there is an important distinction between them data are individual pieces of factual information recorded and used for the purpose of analysis.
It is the raw information from which statistics are created. Statistics are the results of data analysis - its interpretation and presentation. In other words some computation has taken place that provides some understanding of what the data means. Statistics are often, though they don’t have to be, presented in the form of a table, chart, or graph. Both statistics and data are frequently used in scholarly research. Statistics are often reported by government agencies - for example, unemployment statistics or educational literacy statistics. Often these types of statistics are referred to as 'statistical data'.
1
Mid Term Assignment
ASSIGNMENT 1- BRIEF I N T R O D U C T I O N
1. Mean The mean is also referred to as the average, and it is the most commonly used among the three measures of central tendency. The mean is obtained by summing and dividing the values by the number of scores. For example, in five households that comprise 5, 2, 1, 3, and 2 children, the mean can be calculated as follows: = (5+2+1+3+2)/5 = 13/5 = 2.6
2. Median The median is used to calculate variables that are measured with ordinal, interval, or ratio scales. It is obtained by arranging the data from the lowest to the highest and then picking the number(s) in the middle. If the total number of data points is an odd number, the median is usually the middle number. If the numbers are even, the median is obtained by summing the two numbers in the middle and dividing them by two to get the mean. Median is mostly used when there are a few data points that are different. 17, 17, 18, 19, 19, 20, 21, 25, 28, 32 The median of the values above is (19+20)/2 = 19.5. Mode The mode is the most occurring number within a data distribution. It shows what number or value is the highest in number or most common in the data distribution. The mode is used for any type of data. For example, let’s take the example of a college class with about 40 students. The students are given a test exam, graded, and then grouped on a scale of 1-5, starting with students with the lowest number of marks. The marks are graded as follows: •Cluster 1: 5 •Cluster 2: 7 •Cluster 3: 13 •Cluster 4: 12 •Cluster 5: 3 Cluster 3 shows the highest number of students and, therefore, the mode is 13. It reveals that out of 40 students, most of the students were graded in cluster 3.
2
Mid Term Assignment
TASK 1 S. NO. Name of Student
A S S I G N M E N T 1
Subject Wise Marks English
Maths Science S.ST. Hindi
Score/ Grand Total
Percentage (%)
Computer Science
Out of 600
Out of 100
1
Paras
94
85
96
92
98
95
560
93%
2
Aisha
74
70
56
63
78
80
421
70%
3
Ambika
60
56
74
66
85
82
423
71%
4
Manav
90
87
88
83
94
89
531
89%
5
Ellora
69
40
33
54
66
70
332
55%
6
Kanishk
58
66
40
38
39
48
289
48%
7
Ram
44
54
59
45
75
65
342
57%
8
Juhi
85
95
75
85
81
78
499
83%
9
Sanya
79
77
72
84
85
77
474
79%
10
Amrit
87
59
60
66
92
91
455
76%
MEAN
74
68.9
65.3
67.6
79.3
77.5
432.6
MAX.
94
95
96
92
98
95
560
MIN.
44
40
33
38
39
48
289
MEDIAN
76.5
68
66
66
83
79
439
PASS
Greater than 50 Less than 50
3
Mid Term Assignment
TASK 2 Sr. NO. Designation Salary Pf (20% of salary)
A S S I G N M E N T
1
Urban Planner 100000
2 3 4
Architect Engineer
Column2
20000
200000 50000
40000 10000
Urban Designer 25000
5000
Mean Median Mode Maxima Minima
Column1 Mean
93750 Mean 18750 38696.1992 Standard Error Standard Error 7739.239842 1 Median 75000 Median 15000
93750 75000 #N/A
Mode Standard Deviation
#N/A 77392.3984 2
Sample Variance 5989583333 Kurtosis
250000 200000
Skewness
150000 100000
Range Minimum Maximum Sum Count
50000 0
1 Salary
Pf (20% of salary)
Expon. (Salary)
Expon. (Pf (20% of salary))
4
0.75765595 5 1.13762436 7 175000 25000 200000 375000 4 0
Mode Standard Deviation Sample Variance
#N/A 15478.47968
239583333.3
Kurtosis
0.757655955
Skewness
1.137624367
Range Minimum Maximum Sum Count
35000 5000 40000 75000 4
Mid Term Assignment
TASK 3 A S S I G N M E N T
Sr. No.
Name of Student
Hindi
Computer Science
1
Paras
94
85
96
92
98
95
560
93%
A
2
Aisha
74
70
56
63
78
80
421
70%
B
3
Ambika
60
56
74
66
85
82
423
71%
B
4
Manav
90
87
88
83
94
89
531
89%
A
5
Ellora
69
40
33
54
66
70
332
55%
FAIL
6
Kanishk
58
66
40
38
39
48
289
48%
FAIL
7
Ram
44
54
59
45
75
65
342
57%
FAIL
8
Juhi
85
95
75
85
81
78
499
83%
A
1
9
Sanya
79
77
72
84
85
77
474
79%
B
10
Amrit
87
59
60
66
92
91
455
76%
B
English Maths Science S.ST.
5
Total Percentage Grade
Mid Term Assignment
A S S I G N M E N T
Row Labels
Sum of sr. no.
Aisha B Ambika B Amrit B Ellora FAIL Juhi A Kanishk FAIL Manav A Paras A Ram FAIL Sanya B Grand Total
2 2 3 3 10 10 5 5 8 8 6 6 4 4 1 1 7 7 9 9 55
Sum of English 74 74 60 60 87 87 69 69 85 85 58 58 90 90 94 94 44 44 79 79 740
Sum of Maths 70 70 56 56 59 59 40 40 95 95 66 66 87 87 85 85 54 54 77 77 689
Sum of Science 56 56 74 74 60 60 33 33 75 75 40 40 88 88 96 96 59 59 72 72 653
Sum of S.ST. 63 63 66 66 66 66 54 54 85 85 38 38 83 83 92 92 45 45 84 84 676
Sum of Hindi 78 78 85 85 92 92 66 66 81 81 39 39 94 94 98 98 75 75 85 85 793
Sum of Computer Science 80 80 82 82 91 91 70 70 78 78 48 48 89 89 95 95 65 65 77 77 775
Sum of Total 421 421 423 423 455 455 332 332 499 499 289 289 531 531 560 560 342 342 474 474 4326
Sum of Percentage 0.701666667 0.701666667 0.705 0.705 0.758333333 0.758333333 0.553333333 0.553333333 0.831666667 0.831666667 0.481666667 0.481666667 0.885 0.885 0.933333333 0.933333333 0.57 0.57 0.79 0.79 7.21
300 200 100 0
1 English
Maths
6
Science
Mid Term Assignment
TASK 4 A S S I G N M E N T 1
Sr. No. First Name Paras 1 Sheena 2 Ellora 3 Parth 4 Lokesh 5 6 7 8 9 10
Last Name Mongia Sharma Ghosh Lekhi Meena
Combine Paras_Mongia Sheena_Sharma Ellora _Ghosh Parth_Lekhi Lokesh_Meena
Length 12 13 13 11 12
Product Quantity Cost Sofa 5 500 Chair 4 440 tv 2 200 ac 7 655 fridge 1 141
Anisha
Sharma
Anisha_Sharma
13
refrigirat or
3
485
NW
Shubhra Diksha Rahul Yashika
Sharma Tiwari Kasat Singhal
Shubhra_Sharma Diksha_Tiwari Rahul_Kasat Yashika_Singhal
14 13 11 15
fan cooler table table fan
2 8 9 1
185 159 946 154
NW NW W W
count if product is working
4
cont if product is TV and working
1
sum if value is > 300 sumifs
Status W NW W NW NW
3026 200
7
Mid Term Assignment
TASK 1 Chi Test - Calculate and Interpret Chi Square in SPSS
A S S I G N M E N T
RELEVANCE : A chi-square (χ2) statistic is a test that measures how a model compares to actual observed data. The data used in calculating a chi-square statistic must be random, raw, mutually exclusive, drawn from independent variables, and drawn from a large enough sample. Chi-square tests are often used in hypothesis testing. The chi-square statistic compares the size of any discrepancies between the expected results and the actual results, given the size of the sample and the number of variables in the relationship. QUESTION : In the sample dataset, respondents were asked their gender and whether they were a cigarette smoker. We have to check the association between the two using Chi-Square Test of Independence (using α = 0.05).
STEPS OF THE TEST : 1.
2.
3. 4.
5.
Click on Analyze -> Descriptive Statistics -> Crosstabs. Drag and drop (at least) one variable into the Row(s) box (ex. Smoking), and (at least) one into the Column(s) box (ex. Gender). Click on Statistics, and select Chi-square. Press Continue, and then OK to do the chi square test. The result will appear in the SPSS output viewer.
HYPOTHESIS : The approach is to test assumed/ observed values in the data to expected values to check the null value’s truthfulness. Null Hypothesis is accepted when the value of alpha is less than 0.05
ASSUMPTIONS
STEP 1
STEP 2
DATASET
Just like any other statistical test, the chi-square test comes with a few assumptions of its own: •
•
The χ2 assumes that the data for the study is obtained through random selection, i.e. they are randomly picked from the population The categories are mutually exclusive i.e. each subject fits in only one category. For e.g.- from our above example – the number of people who • lunched in your restaurant on Monday • can’t be filled in the Tuesday category •
The data should be in the form of frequencies or counts of a particular category and not in percentages The data should not consist of paired samples or groups or we can say the observations should be independent of each other When more than 20% of the expected frequencies have a value of less than 5 then Chi-square cannot be used. To tackle this problem: Either one should combine the categories only if it is relevant or obtain more data
OUTPUTS OF THE TEST
2
RESULT Since, the value of α is greater than 0.05. So, hypothesis is not rejected. There is a relationship between smoking & gender variable.
8
End Term Assignment
TASK 2
One Sample T Test A S S I G N M E N T 2
Calculate and Interpret
RELEVANCE : The one-sample t-test is a statistical hypothesis test used to determine whether an unknown data mean is different from a specific value. There are three types of t-tests we can perform based on the data at hand: • One sample t-test • Independent two-sample t-test • Paired sample t-test
Hence, we can perform a onesample t-test. Here’s the formula to calculate this:
t = t-statistic m = mean of the group µ = theoretical value or population mean s = standard deviation of the group n = group size or sample size
QUESTION : There is a tire making company which claims that their tires can give an average of 35-45km. Make an Analysis to check will the tire at 40k will be workable or not.
STEPS OF THE TEST : 1. Click on Analyze -> Compare Means -> One-Sample T Test 2. Drag and drop the variable you want to test against the mean into the Test Variable(s) box. 3. Specify a mean in the Test Value box 4. Click OK 5. Results will appear in the SPSS output viewer
HYPOTHESIS : If H0 > 0.05, then the hypothesis is correct
DATASET
OUTPUTS OF THE TEST
Normality Test is not significantly related.
The value of α is > 0.05, hence the hypothesis is accepted and null is verified.
ASSUMPTIONS There are certain assumptions we need to heed before performing a ttest: 1. The data should follow a continuous or ordinal scale (the IQ test scores of students, for example) 2. The observations in the data should be randomly selected 3. The data should resemble a bell-shaped curve when we plot it, i.e., it should be normally distributed. 4. Large sample size should be taken for the data to approach a normal distribution (although t-test is essential for small samples as their distributions are non-normal)
RESULT
9
End Term Assignment
TASK 3 INDEPENDENT T Test
A S S I G N M E N T 2
-
Calculate and Interpret
RELEVANCE : The Independent Samples t Test compares the means of two independent groups in order to determine whether there is statistical evidence that the associated population means are significantly different. The Independent Samples t Test is a parametric test. The Independent Samples t Test is commonly used to test the following: • Statistical differences between the means of two groups • Statistical differences between the means of two interventions • Statistical differences between the means of two change scores
QUESTION : A construction company plans to construct property at Delhi and Mumbai, but the company needs to know that if the property rates are same in both the cities or different.
STEPS OF THE TEST : 1. Click on Analyze -> Compare Means -> Independent-Samples T Test 2. Drag and drop the dependent variable into the Test Variable(s) box, and the grouping variable into the Grouping Variable box 3. Click on Define Groups, and input the values that define each of the groups that make up the grouping variable (i.e., the coded value for Group 1 and the coded value for Group 2) 4. Press Continue, and then click on OK to run the test 5. The result will appear in the SPSS data viewer
ASSUMPTIONS 1.
2.
3.
The dependent variables should be measured on a continuous scale (either interval or ratio). There should be two dependent variables present which are measured from independent (non-related) groups. There are no outliers present in the variables.
HYPOTHESIS : If H0 > 0.05, Paying Capacity of Delhi and Mumbai people are same
DATASET (179 samples)
OUTPUTS OF THE TEST Independent T Test Total Samples = 179 Delhi – 1 (80 people) Mumbai – 2 (90 people)
RESULT
10
City Amount
Statistic
df
Shapiro-Wilk Sig.
Statistic
df
Sig.
Delhi
.219
80
.000
.864
80
.000
Mumbai
.153
99
.000
.824
99
.000
Group Statistics City Amount
N
Mean
Std. Deviation
Std. Error Mean
Delhi
80
4587500.00
842596.199
94205.119
Mumbai
99
4632727.27
1031639.094
103683.630
Independent Samples Test
The value of α is > 0.05, the hypothesis is verified. 4. The dependent variables should be normally distributed. 5. The dependent variables should have homogeneity of variances. In other words, their standard deviations need to be approximately the same. This can be investigated with the Levene’s Test for Equality of Variances.
Tests of Normality Kolmogorov-Smirnova
Levene's Test for
t-test for
Equality of
Equality of
Variances
Means
F Amo Equal unt
.234
Sig.
t
df
.629 -.316
177
-.323
176.
variances assumed Equal variances not
975
assumed
End Term Assignment