library(googlesheets4)
library(ggplot2); library(dplyr)
#url <- "URL TO SHEET HERE" # hiddenSex Discrimination Case Study
Load libraries and specify URL to google sheets (hidden)
Import data and visualize
sim.data <- read_sheet(url) %>%
select(mfdiff = `(Male - female) difference`) %>%
mutate(pt_estimate = round(mfdiff, 2))
ggplot(sim.data, aes(x = mfdiff)) +
geom_dotplot() + theme_bw() +
geom_vline(xintercept = .292, color = "blue") +
xlab("point estimate")
Describe the distribution of this graph. What does it seem to be centered around?
The distribution of the point estimate (\(p_m-p_f\)) is bell shaped and symmetric around 0, with an outlier up around 0.4. All simulated differences except one are below the 0.292 mark.
In what percent of simulations did we observe a difference of at least 29.2% (0.292)?
table(sim.data$pt_estimate)
-0.21 -0.12 -0.04 0.04 0.21 0.29
1 1 6 4 1 1
table(abs(sim.data$mfdiff) >= .29)
FALSE TRUE
13 1
In our experiment we observed 1/14 of the simulations resulted in a point estimate as extreme as the one in the original study. We determined that there was a 7.1% chance of obtaining a sample where \(\geq\) 29% more male candidates than female candidates get promoted under the null hypothesis of this event happening by chance (yes I rounded .292 to .29 here).
We determined that this proportion is small enough to conclude that the data provide evidence of sex discrimination against female candidates by the male supervisors.