{"id":756,"date":"2009-07-10T15:54:43","date_gmt":"2009-07-10T19:54:43","guid":{"rendered":"http:\/\/mat.tepper.cmu.edu\/blog\/?p=756"},"modified":"2009-07-10T15:54:43","modified_gmt":"2009-07-10T19:54:43","slug":"the-perils-of-statistical-significance","status":"publish","type":"post","link":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/2009\/07\/10\/the-perils-of-statistical-significance\/","title":{"rendered":"The Perils of &#8220;Statistical Significance&#8221;"},"content":{"rendered":"<p>As someone who teaches data mining, which I see as part of operations research, I often talk about what sort of results are worth changing decisions over.\u00a0 Statistical significance is not the same as changing decisions.\u00a0 For instance, knowing that a rare event is 3 times more likely to occur under certain circumstances might be statistically significant, but is not &#8220;significant&#8221; in the broader sense if your optimal decision doesn&#8217;t change.\u00a0 In fact, with a large enough sample, you can get &#8220;statistical significance&#8221; on very small differences, differences that are far too small for you to change decisions over.\u00a0 &#8220;Statistically different&#8221; might be necessary (even that is problematical) but is by no means sufficient when it comes to decision making.<\/p>\n<p>Finding statistically significant differences is a tricky business.\u00a0 <a href=\"http:\/\/www.americanscientist.org\/\"><em>American Scientist<\/em><\/a> in its July-August, 2009 issue has a <a href=\"http:\/\/www.americanscientist.org\/issues\/feature\/2009\/4\/of-beauty-sex-and-power\">devastating article<\/a> by Andrew Gelman and David Weakliem regarding the research of <a href=\"http:\/\/www2.lse.ac.uk\/researchAndExpertise\/Experts\/s.kanazawa@lse.ac.uk\">Satoshi Kanazawa<\/a> of the London School of Economics.\u00a0 I highly recommend the article (available at one of the <a href=\"http:\/\/www.stat.columbia.edu\/~cook\/movabletype\/archives\/2009\/06\/of_beauty_sex_a.html\">authors&#8217; sites<\/a>, along with a discussion of the article,\u00a0 and I definitely recommend buying the magazine, or subscribing:\u00a0 it is my favorite science magazine) as a lesson for what happens when you get the statistics wrong.\u00a0 You can check out the whole article, but perhaps you can get the message from the following graph (from the <em>American Scientist<\/em> article):<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-757\" title=\"gelman analysis of Kanazawa\" src=\"https:\/\/fac-mtrick02.tepper.cmu.edu\/blog\/wp-content\/uploads\/2009\/07\/gelman.jpg\" alt=\"gelman analysis of Kanazawa\" width=\"500\" height=\"346\" \/><\/p>\n<p>The issue is whether attractive parents tend to have more daughters or sons.\u00a0 The dots represent the data:\u00a0 parents have been rated on a scale of 1 (ugly) to 5 (beautiful) and the y-axis is the fraction of their children who are girls.\u00a0 There are 2972 respondents\u00a0 in this data.\u00a0 Based on the Gelman\/Weakliem discussion, Kanazawa (in the respected journal <em>Journal of Theoretical Biology<\/em>) concluded that, yes, attractive parents have more daughters.\u00a0 I have not read the Kanazawa article, but the title &#8220;Beautiful Parents Have More Daughters&#8221; doesn&#8217;t leave a lot of wiggle room (Journal of Theoretical Biology, 244: 133-140 (2007)).<\/p>\n<p>Now, looking at the data suggests certain problems with that conclusion.\u00a0 In particular, it seems unreasonable on its face.\u00a0 With the ugliest group having 50-50 daughters\/sons, it is really going to be hard to find a trend here.\u00a0 But if you group 1-4 together and compare it to 5, then you can get statistical significance.\u00a0 But this is statistically significant only if you ignore the possibility to group 1 versus 2-5, 1-2 versus 3-5, and 1-3 versus 4-5.\u00a0 Since all of these could result in a paper with the title &#8220;Beautiful Parents Have More Daughters&#8221;, you really should include those in your test of statistical significance.\u00a0 Or, better yet, you could just look at that data and say &#8220;I do not trust any test of statistical significance that shows significance in this data&#8221;.\u00a0 And, I think you would be right.\u00a0 The curved lines of Gelman\/Weakliem in the diagram above are the results of a better test on the whole data (and suggest there is no statistically significant difference).<\/p>\n<p>The <em>American Scientist<\/em> article makes a much stronger argument regarding this research.<\/p>\n<p>At the recent <a href=\"http:\/\/www.euro-2009.de\">EURO<\/a> conference, I attended a talk on an aspects of sports scheduling where the author put up a graph and said, roughly, &#8220;I have not yet done a statistical test, but it doesn&#8217;t look to be a big effect&#8221;.\u00a0 I (impolitely), blurted out &#8220;I wouldn&#8217;t trust any statistical test that said this was a statistically significant effect&#8221;.\u00a0 And I think I would be right.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>As someone who teaches data mining, which I see as part of operations research, I often talk about what sort of results are worth changing decisions over.\u00a0 Statistical significance is not the same as changing decisions.\u00a0 For instance, knowing that a rare event is 3 times more likely to occur under certain circumstances might be &hellip; <a href=\"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/2009\/07\/10\/the-perils-of-statistical-significance\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;The Perils of &#8220;Statistical Significance&#8221;&#8221;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[15,46],"tags":[],"class_list":["post-756","post","type-post","status-publish","format-standard","hentry","category-data-mining","category-research"],"jetpack_featured_media_url":"","_links":{"self":[{"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/posts\/756","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/comments?post=756"}],"version-history":[{"count":0,"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/posts\/756\/revisions"}],"wp:attachment":[{"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/media?parent=756"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/categories?post=756"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/tags?post=756"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}