{"id":687,"date":"2009-05-18T15:05:24","date_gmt":"2009-05-18T19:05:24","guid":{"rendered":"http:\/\/mat.tepper.cmu.edu\/blog\/?p=687"},"modified":"2009-05-18T15:05:24","modified_gmt":"2009-05-18T19:05:24","slug":"data-mining-competition-from-fico-and-ucsd","status":"publish","type":"post","link":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/2009\/05\/18\/data-mining-competition-from-fico-and-ucsd\/","title":{"rendered":"Data Mining Competition from FICO and UCSD"},"content":{"rendered":"<p>I am a sucker for competitions.\u00a0 I have run a few in the past, and I see my page on the <a href=\"http:\/\/mat.tepper.cmu.edu\/TOURN\">Traveling Tournament Problem<\/a> as an indefinite length computational competition.\u00a0\u00a0\u00a0 Data Mining naturally leads to competitions:\u00a0 there are so many alternative techniques out there and little idea of what might work well or poorly on a particular data set.\u00a0 The preeminent challenge of this type is the <a href=\"http:\/\/www.netflixprize.com\/\">Netflix Prize<\/a>, where the goal is to better predict customer movie ratings, and win a million dollars in doing so.\u00a0 I have <a href=\"http:\/\/mat.tepper.cmu.edu\/blog\/?p=174\">written before<\/a> about the lessons to be learned from this particular challenge (in short:\u00a0 while it might be a nice exercise, it is pretty clear that the improvements given by the algorithm would have little noticeable effect on the customer experience).<\/p>\n<p><a href=\"http:\/\/www.fico.com\">FICO<\/a> (formerly known as FairIsaac) has sponsored a data mining competition with the University of California San Diego for a number of years.\u00a0 The competition is open to all students (and postdocs) and have just announced the 2009 competition.\u00a0 The <a href=\"http:\/\/mill.ucsd.edu\/\">website for the competition is now open<\/a>, with a finish date for the competition of July 15, 2009.<\/p>\n<p>The data involves detecting anomalous e-commerce transactions and come in &#8220;easy&#8221; and &#8220;hard&#8221; versions.\u00a0 I have spent a couple of minutes with the data and it is quite interesting to work with.<\/p>\n<p>I do have one complaint with this sort of data mining.\u00a0 In my data mining class, I stress that you can do much better data mining if you understand the business context.\u00a0 This understanding need not be overly deep, but it is hard to analyze data that is simply given as &#8220;field1&#8221;, &#8220;field2&#8221;, and so on.\u00a0 For problems where creating new fields is important (say, aggregating ten types of insurance policies into one new field giving number of insurance policies purchased), if you don&#8217;t understand what the data means, it is impossible to generate appropriate new fields.\u00a0 The data set in this competition has had its fields anonymized so strongly that finding any creative new fields will be more a matter of luck than anything else.<\/p>\n<p>Despite this caveat, I think the competition is a great chance for students to show off what they have learned or developed.\u00a0 It would be particularly nice for an operations research approach to do well.\u00a0 And it doesn&#8217;t last forever like the Traveling Tournament Problem or, it seems, the Netflix Prize.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>I am a sucker for competitions.\u00a0 I have run a few in the past, and I see my page on the Traveling Tournament Problem as an indefinite length computational competition.\u00a0\u00a0\u00a0 Data Mining naturally leads to competitions:\u00a0 there are so many alternative techniques out there and little idea of what might work well or poorly on &hellip; <a href=\"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/2009\/05\/18\/data-mining-competition-from-fico-and-ucsd\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Data Mining Competition from FICO and UCSD&#8221;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[10,15],"tags":[],"class_list":["post-687","post","type-post","status-publish","format-standard","hentry","category-challenges","category-data-mining"],"jetpack_featured_media_url":"","_links":{"self":[{"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/posts\/687","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/comments?post=687"}],"version-history":[{"count":0,"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/posts\/687\/revisions"}],"wp:attachment":[{"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/media?parent=687"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/categories?post=687"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/mat.tepper.cmu.edu\/blog\/index.php\/wp-json\/wp\/v2\/tags?post=687"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}