Wednesday, January 2, 2019

Statistics

Statistics is a field of study concerned with (1) the collection, organization summarization, and analysis of data; (2) the drawing of inference about a body of data when only a part of data is observed.

A descriptive measure computed from the data of a sample is called a statistic


A descriptive measure computed from the data of a population is called a parameter
Variable: A characteristic that takes different values in different persons, places or things.

Quantitative variable: that can be measured in the usual sense. Measurement convey information about amount.

Qualitative variable: measurement consist of categorization. Measurement convey information regarding attribute.

Random variable: when the values arise as a result of chance factor, so they cannot be predicted in advance 

Discrete variable: is characterized by gaps or interruptions in the values that it can assume 

Continuous variable: doesn’t possess the gaps or interruptions characteristic of a discrete variable

Population: largest collection of entities for which we have an interest at a particular time

Sample: a part of the population that we took for studying 

Measurement: assignment of numbers to objects or events according to as set of rules. Measurement may be carried out under different sets of rules

Measurement Scale: 

Nominal scale: naming the observations or classifying them into various mutually exclusive and collective exhaustive categories

Ordinal scale: when observations are not only from different categories but also can be ranked according to some criteria 

Interval scale: in addition to ordering the measurement we can also know the distance between the two measurements
Interval scale unlike the nominal and ordinal scales is a truly quantitative scale

Ratio scale: highest level of measurement. Equality of ratios as well as equality of the intervals may be determined. 

Fundamental to the ratio scale is true zero point


Simple random sample: If a sample of size “n” is drawn from a population of size “N” in such a way that every possible sample of size “n” has the same chance of being selected, the sample is called simple random sampling
As a rule, in practice, sampling is always done without replacement.

Systematic sampling: first we calculate the total number required for the sample, a random number table is then used to give a starting number (x). A second number determined by the sample size is selected to define the sampling interval (k). Now we select individuals in this way
x, x+k, x+2k, x+3k, …….


Stratified random sampling: population is stratified into strata. And a random sampling is taken in each strata.
Creative Commons License
PSM / COMMUNITY MEDICINE by Dr Abhishek Jaiswal is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Based on a work at learnpsm@blogspot.com.
Permissions beyond the scope of this license may be available at jaiswal.fph@gmail.com.

Box and whisker plots (Boxplot)


Box and whisker plots (Boxplot):

Represents the variable of interest on horizontal axis
A box is drawn such a way that left end of box align with Q1, and the right end align with Q3.
Divide the box into two parts by a vertical line that aligns with the median Q2
Draw a horizontal line called a whisker from the left end of the box to the point that align with the smallest measurement of the data set
Draw another horizontal line or whisker from the right end of the box to the point that align with the largest measurement of the data set






Outliers: it is an observation whose value, “x”, either exceeds the value of the third quartile by a magnitude greater than 1.5(IQR) or is less than the value of first quartile by a magnitude greater than 1.5(IQR).

That is

{Q1- 1.5(IQR)} > x > {Q3 + 1.5(IQR)}

Creative Commons License
PSM / COMMUNITY MEDICINE by Dr Abhishek Jaiswal is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Based on a work at learnpsm@blogspot.com.
Permissions beyond the scope of this license may be available at jaiswal.fph@gmail.com.

Stem and leaf


Stem and leaf: we partition each measurement into two parts. The first part is called the stem, the second is called the leaf. The stem consists of one or more of the initial digits of the measurement, the leaf is composed of one or more of the remaining digits. All the partitioned numbers are shown together in a single display; the stems form an ordered column with the smallest stem at the top and the largest at the bottom. We include in the stem column all stems within the range of the data even when a measurement with that stem is not in the data set. The rows of display contain the leaves, ordered and listed to the right of their respective stems. When leaves consist of more than one digit, all digits after the first may be deleted. Decimal when present in data is omitted in the stem and leaf display. The stems are separated from their leaves by a vertical line.

An advantage of it over histogram is that it preserves the information contained in the individual measurements.
Are most effective with relatively small data sets

Creative Commons License
PSM / COMMUNITY MEDICINE by Dr Abhishek Jaiswal is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Based on a work at learnpsm@blogspot.com.
Permissions beyond the scope of this license may be available at jaiswal.fph@gmail.com.

Meta-analysis


Meta-analysis:

Effect size: The effect size is a value which reflects the magnitude of the treatment effect or the strength of a relationship between two variables, is the unit of currency in meta-analysis. (Black square)
Precision: It is the C.I. of the effect-size.
Study weight: It is the weight assigned to each study. The weight assigned is dependent on the precision. (The larger the square the larger the study weight)
Summary effect: It is the weighted mean of the individual effects. The mechanism used to assign the weights depends on our assumptions about the distribution of effect sizes from which the studies were sampled. Under the fixed-effect model, the assumption is that all the studies in the analysis share the same true effect size. The summary effect then is the estimate of this common effect size. Under the random-effect model, the assumption is true effect size varies from study to study. The summary effect here will be the mean of the distribution of effect sizes. (Diamond)
Precision: The location of the diamond represents the effect size. Its width reflects the precision of the estimate.

In fixed effect model: the weight of individual study is reciprocal of that study’s variance




In Random-effect model: We assume that the true effect is normally distributed


To be continued...

Creative Commons License
PSM / COMMUNITY MEDICINE by Dr Abhishek Jaiswal is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Based on a work at learnpsm@blogspot.com.
Permissions beyond the scope of this license may be available at jaiswal.fph@gmail.com.

Sensitivity and Specificity


Validity (accuracy): Extent to which a test measures what it is supposed to measure.

Sensitivity:      1. Ability of test to correctly classify an individual as diseased.
                        2. Probability of being test positive when disease is present.


D+
D-
T+
A
B
T-
C
D



SnNOUT: Highly sensitive test if negative rules out the disease

Specificity:      1. Ability of test to correctly classify an individual as disease free.
                        2. Probability of being test negative when disease is absent.

                                                                                                



SpPIN: Highly specific test if positive rules in the disease.

PPV: Positive predictive value:
1.     % of patients with positive test who actually have the disease
2.     Probability of patient having disease when test is positive








NPV: Negative predictive value:
1.     % of patients having disease when test is positive
2.     probability of patient having disease when test is positive






Bayes Theorem:





PPV: Highly dependent on prevalence of disease






Parallel testing:

A-test or B-test: (A, B) sensitivity or specificity
Combined sensitivity: Sn= A+B-AB
Combined specificity: Sp=A*B
Sensitivity will increase and specificity will decrease

Series testing:

A-test or B-test: (A, B) sensitivity or specificity
Combined sensitivity: Sn= A*B
Combined specificity: Sp=A+B-AB
Sensitivity will decrease and specificity will increase

Creative Commons License
PSM / COMMUNITY MEDICINE by Dr Abhishek Jaiswal is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Based on a work at learnpsm@blogspot.com.
Permissions beyond the scope of this license may be available at jaiswal.fph@gmail.com.

Mantel Haenszel

Mantel Haenszel method is one of the method to control for confounders. It gives a single summary measure of association which provides a weighted average of RR or OR across different strata of confounding factors.
To calculates in this method we first have to divide the original two by two table by different strata of confounding variable and then we calculate the weighted average of RR or OR.
formula
                                          outcome (O)
                                             +      -   
RISK FACTOR (E)     +      a.    b.       a+b
                                     -       c.    d.       c+d
         
                                             a+c.  b+d.  
RR =   (a/(a+b)) ÷ (c/(c+d))     =    a(c+d)÷ c(a+b)
OR =    a/b.  ÷   c/d.        =   ad/bc
RR (mh)     = summation (a(c+d)÷n) ÷ summation (c(a+b) ÷n)
OR (mh)     = summation (ad/n) ÷summation (bc/n)
summation is sigma, that is sum of all the values in the different two by two tables
Creative Commons License
PSM / COMMUNITY MEDICINE by Dr Abhishek Jaiswal is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Based on a work at learnpsm@blogspot.com.
Permissions beyond the scope of this license may be available at jaiswal.fph@gmail.com.

Skewness and Kurtosis

Skewness: If a histogram/frequency polygon of a distribution is asymmetric, the distribution is said to be skewed. 

If the distribution is not symmetric because its graph extends further to the right than to the left, that is, if it has a long tail to the right, we say that the distribution is skewed to the right or it is positively skewed. 
A distribution will be skewed to the right, or positively skewed, if its mean is greater than its mode.

If the distribution is not symmetric because its graph extends further to the left than to the right, that is, if it has a long tail to the left, we say that the distribution is skewed to the left or it is negatively skewed. 
A distribution will be skewed to the left, or negatively skewed, if its mean is less than its mode.






Skewness >0 indicates positive skewness

                 <0 indicates negative skewness



Kurtosis: it is a measure of the degree to which the distribution is peaked or flat in comparison to a normal distribution whose graph is characterized by a bell shaped appearance.

Platykurtic: the graph exhibits a flattened appearance 
Mesokurtic: normal, bell shaped graph
Leptokurtic: the graph exhibits a more peaked appearance 





Platykurtic kurtosis <0
Mesokurtic kurtosis =0
Leptokurtic kurtosis >0

Kurtosis: it is a measure of the degree to which the distribution is peaked or flat in comparison to a normal distribution whose graph is characterized by a bell shaped appearance.

Platykurtic: the graph exhibits a flattened appearance
Mesokurtic: normal, bell shaped graph
Leptokurtic: the graph exhibits a more peaked appearance

Creative Commons License
PSM / COMMUNITY MEDICINE by Dr Abhishek Jaiswal is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Based on a work at learnpsm@blogspot.com.
Permissions beyond the scope of this license may be available at jaiswal.fph@gmail.com.

Featured Post

Sample Size Calculator

Advanced EpiCalc Pro - Comprehensive Sample Size Calculator ...

Popular Posts