Sex Discrimination Case Study

Load libraries and specify URL to google sheets (hidden)

library(googlesheets4)
library(ggplot2); library(dplyr)
#url <- "URL TO SHEET HERE" # hidden

Import data and visualize

sim.data <- read_sheet(url) %>% 
  select(mfdiff = `(Male - female) difference`) %>%
  mutate(pt_estimate = round(mfdiff, 2))

ggplot(sim.data, aes(x = mfdiff)) + 
  geom_dotplot() + theme_bw() + 
  geom_vline(xintercept = .292, color = "blue") +
  xlab("point estimate")

Tip

Describe the distribution of this graph. What does it seem to be centered around?

The distribution of the point estimate (\(p_m-p_f\)) is bell shaped and symmetric around 0, with an outlier up around 0.4. All simulated differences except one are below the 0.292 mark.

Tip

In what percent of simulations did we observe a difference of at least 29.2% (0.292)?

table(sim.data$pt_estimate) 

-0.21 -0.12 -0.04  0.04  0.21  0.29 
    1     1     6     4     1     1 
table(abs(sim.data$mfdiff) >= .29)

FALSE  TRUE 
   13     1 

In our experiment we observed 1/14 of the simulations resulted in a point estimate as extreme as the one in the original study. We determined that there was a 7.1% chance of obtaining a sample where \(\geq\) 29% more male candidates than female candidates get promoted under the null hypothesis of this event happening by chance (yes I rounded .292 to .29 here).

We determined that this proportion is small enough to conclude that the data provide evidence of sex discrimination against female candidates by the male supervisors.