Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Data was from random volunteers, and at that moment in time over 10% of the U.S. employee base. I say size was statistically significant not quality of data because there were senior engineers who had been with the company for 15+ years and would skew the datasets. We had lots of demographic data too. Can you guess which race and gender won?


>I say size was statistically significant not quality of data

Size is not necessarily significant unless is a random selected sample (and I'm not sure volunteers would count as random).

Let's say you have a population of 100 and there are two groups (A-80% and B- 20%). Let's say that individuals in group A are more keen to voluteer themselves. You could have a sample of 50% of the population and still not have a representative sample of their entire population (ie. 49 from A and 1 from B), so the sample size didn't matter to much in that case.


> Size is not necessarily significant unless is a random selected sample (and I'm not sure volunteers would count as random).

I would say that most surveys/studies are done on a voluntary basis.


There is also a difference in e.g. stopping random people on the street and asking if they are willing to participate, or posting an ad online and waiting for people to call, even though both are voluntary.


>I would say that most surveys/studies are done on a voluntary basis.

Also most surveys/studies carefully say that they are not necessarily representative.


Yes, a disclaimer over the validity of the results would be nice ;). But to be fair, most surveys are taken rather seriously, especially by the common populace and the main stream media, ignoring the lack of absoluteness in the results.


There is no such thing as a random volunteer.


True, but there are ways to mitigate that. YouGov tries to create a representative sample out of a batch of volunteers.

https://en.wikipedia.org/wiki/YouGov


I was very curious as to how they would attempt to mitigate that, but your link doesn't actually provide any information about it.

I think I would somewhat cautiously state that it isn't something that is possible to mitigate. There are simply two cohorts: people who are willing to take online polls and people who are not. It is inherently impossible to gather this kind of data on the latter.


I would be careful with defining winners and losers. How does pay distribution inside any given group look like? What about the age distribution? Health? Overtime policy? Who assigns work and the degree of responsibility attached to each assignment?

For example, if we define that white males get the highest page in average, we could also see in the same statistics that middle aged white males who suffer from health issues and can't compete in their group in overtime get less pay compared to someone with the same health and age group but of different race and gender. Are white males then winners?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: