Search This Blog

Saturday, April 18, 2015

Women's Healthcare Prediction


The Women's Healthcare Prediction competition was held on DrivenData from Feb-2015 to Apr-2015.

Objective
The challenge was to predict which healthcare services (like household, pregnancy, family, medical, etc.) were opted by women. Essentially, it was a multi-label, multi-class classification.

Data
The train data consisted of ~14600 rows (or women) along with various numeric and categorical variables and which of the 14 services were opted by them. Each woman could've opted for more than one service.

The test data consisted of ~3600 rows (or women) for which we had to predict which all services would they have opted for.

Approach
There were 1300+ variables, so my general approach was to do some form of FS along with an ensemble of classification models.

I started off with the usual suspects and found the tree-based models performing better than the linear models. None of the other models came even close to the accuracy received using XGBoost or RandomForest.

I tried multiple ways of doing feature selection and reducing the dimension, but they didn't improve the results significantly.

Once I exhausted all ideas, I used the brute-force approach to optimize my model performance by tweaking the parameters of each of the 14 individual labels.

Model
My final model was an ensemble of XGBoost and RandomForest with some standard data cleaning and FS. I optimized the parameters for each of the 14 labels, but that gave very minor improvement.

Results
I stood 11th on the public LB out of 104 teams. Just missed the Top-10 and also the Top-10% !
My model achieved logloss of 0.2588 while the topper was 0.2539.

View Complete Results

Views
This is the first competition where I really struggled for a long time. Tried lots of ideas, but nothing seemed to work. Ensembles hardly gave any improvement and I was literally stuck during the last 2 weeks.

The public/private LB split seemed excellent with the ranks remaining almost the same. Even the CV and LB scores moved in the same direction.

Feels like I missed out here, but it only motivates me to come harder next time. This was my first competition on DrivenData, and I'm hoping there are better ones to come soon!

Wednesday, April 1, 2015

Unlucky 13


I authored a Sudoku contest Unlucky 13 on LMI. It was held from 1st - 6th April, 2015 and consists of 13 sudokus to be solved in 65 minutes.

View Championship Page

Download Instruction Booklet
Download Puzzle Booklet
Password is LuckyYou

View Forum

View Results

"13 is my favourite number and I created this themed test in late-2014 during some easy days at work. Incidently, this is also the 13th test I'm authoring at LMI. Lot of special moments and memories along the way... and I hope players enjoy this set and make it a success!"

Congrats to Jan Zverina, Hideaki Jo and Jakub Ondrousek for the top-3 overall players.
Congrats to Prakhar Gupta, Kishore Kumar and Rishi Puri for the top-3 Indian players.

"Good artists copy, great artists steal" - Pablo Picasso.

Few months back, a couple of my friends created some sudoku variants and asked me to test solve them. It was their first try at creating sudokus and they did quite a decent job, since they were all unique. Only problem was, all the variants could be solved like Classics without having to use the variant rule and I had a hearty laugh while solving them. For example, there was an Odd Even Sudoku with 42 givens... a Non-Consecutive Sudoku with 33 givens... etc. :-)

That's how the idea was formed for this test. If I enjoyed it so much, maybe other solvers would enjoy it too, in its own humorous way. It was subtle April 'fooling', unlike last year's total surprise (which was awesome in its own way). Thanks to Deb Mohanty and Prasanna Seshadri for test-solving and other inputs and 'contributions'. I'm happy many players were able to complete the test and get the bonus, it was intentionally left longer and the difficulty such that a large portion of solvers would finish.

Thanks for all the messages, and hope to see some more exciting Sudoku solving in the months to come! :-)

Friday, March 20, 2015

Hotel Demand Forecasting


The Hotel Demand Forecasting competition was held on CrowdAnalytix in Feb, 2015.

Objective
The objective was to build a forecasting model to predict the demand in hotel using historic inquiries.

Data
The train data consisted of historic inquiries (reservation, denial, regret) for five different hotels in 2011, 2012, 2013.

The test data was predicting the demand for the hotels in 2014.

Approach
I found the historic demand having strong weekly trends (like you would intuitively expect) and a naive submission of using previous year's demand for the same weekday (eg: using previous year's Friday demand to predict this year's Friday demand in the corresponding week) gave a very good score. So, I decided to play with and optimize the historic averages. I ended up using two different versions of historic averages and made it to the top-5 without a sophisticated model. I wouldn't be surprised if the other toppers used similar ideas.

The first method of historic averages was using a weighted average of the previous three years' demand on the same weekday of the week.

The second method of historic averages was aggregating the weekly demand of the previous three years in the corresponding week and splitting it based on the historic demand proportion for each weekday, which was calculated separately for each quarter.

Model
The final model was an average of the two historic averages, along with some smoothing. The smoothing was redistributing the predictions across three weeks (the week before, the current week and the week after) using a weighted average. This smoothing gave me the biggest jump in my LB score, which pushed me into the top-5.

Code
View on Github

Results
I stood 5th on the public LB out of 45 teams. My model achieved MAPE score of ~ 0.26 and the best was ~ 0.25

The top-5 models were evaluated further and I still stood 5th :-|

Views
This is my second competition on CrowdAnalytix (after Exacerbation) and I'm glad I could finish in the top-5 in both of them. Though these are not as popular as the ones on Kaggle, I enjoyed exploring this forecasting model especially since it was one without any features.

Check out My Best CrowdAnalytix Performances

Tuesday, February 24, 2015

Avazu Click-Through Rate Prediction



The Avazu Click-Through Rate Prediction competition was held on Kaggle from Nov-2014 to Feb-2015.

Objective
Click-through rate is a very important measure for performance of ads and the challenge was to predict how likely an ad will be clicked.

Data
Train data consisted of ~ 40 million ads (which is just 10 days of Avazu data!) along with a label indicating whether they were clicked or not. The variables were about the website/app where the ad appeared, some features of the ad (like size, position, etc.), demographics of the user to whom the ad was shown and some anonymous variables.
The test data consisted of ~ 6 million ads (11th day of Avazu data).

Approach/Model
This is the largest data set I've worked with till date and 40 million rows of data meant memory issues right from the start.

Wait a minute. What about that awesome online-lr code? Of course... that's the same beauty I used for the Tradeshift competition and its the same one I used for this competition too. Well, isn't it just fabulous?

I started off playing around with the parameters of the code and adding interaction variables and generating some features. Some of the anonymous variables were decoded (by some Kagglers) and I tried using them more smartly.

There were massive number of participants and after 2-3 weeks, I was ranked in the top-20 with 600-700 teams. I had some work assignment for which I travelled to US, and wasn't sure if I would have time to try out new ideas, so I decided not to pursue it further.

Not much to share here, no particularly nice model ideas, but I still managed to secure 79th place out of a whopping 1604 teams scoring 0.3908 / 0.3889 using the logloss metric.

Views
It was a challenge to work with this data, and not having access to much RAM, it is all the more tricky. Thanks again to pypy and tinrtgu for the online-lr code and I'm glad I still made it into the top-10%

Congrats to 4-Idiots, Owen and Random Walker for the top-3 spots. What can you say about Owen? Leading the overall Kaggle rankings with more than double the points over 2nd place David Thaler. Some feat that it!

And for me, moved to 185th in overall rankings. The race is on to finish in the top-100 (or top-50) by end of this year.

Check out My Best Kaggle Performances

Saturday, January 10, 2015

Exacerbation Prediction


The Exacerbation Prediction competition was held on CrowdAnalytix in Nov-Dec, 2014.

Objective
Respiratory diseases (asthma, cystic fibrosis, smoking diseases, etc.) are one of the leading causes of deaths globally. As the condition of patients deteriorate, they experience 'exacerbations', which is sudden worsening of symptoms, requiring immediate emergency and medical attention.

The objective was to build a predictive model using medical and genetic data which predicts beforehand which patients will experience 'exacerbation' so that they can be provided appropriate medical treatment to prevent/control it.

Data
The train data consisted of ~ 4000 patients and 1300 variables along with the true labels of whether they experienced Exacerbation or not.
The test data consisted of ~ 2000 patients for which we had to predict the probability of Exacerbation.

Approach
My main idea was to build 2-3 strong classifiers and then build an ensemble with them. With 1300 variables, variable selection / dimension reduction became a must.

I tried tree-based models like Random Forest, GBM, Extra Trees, XG-Boost, etc., regression based models like Logistic Regression, Ridge Regression, etc., and some others like k-NearestNeighbours, SVM, NaiveBayes, etc.

RF and XGB gave the best results while LR and k-NN were decent. I explored and optimized these. After some tuning, XGB and LR gave much better scores, and k-NN didn't add any improvement.

Model
My final model was a weighted average of XG-Boost and Logistic Regression.

The XG-Boost was built on 150 variables, which were selected based on the variable importance of some sample tree models.

The Logistic Regression was built on the top-50 Principal Components.

Results
I stood 4th on the Public LB out of 101 teams, 1st on the Private LB, and finally 2nd on the Private Evaluation. I'm not sharing the scores (since they are not public), but my models achieved AUC scores of ~ 0.845

So, I stood 2nd! This is the first time I've got a ranking with some prize money! Yay!

Views
When I started this competition, I was looking at all numeric features of anonymized variables. I wasn't sure how much I could squeeze out from the data, but I put in a lot of time and effort and found some wonderful ideas in the process.

I think my model was a very robust and competitive one, and I was surprised it scored so consistently across multiple test sets.

Overall, it was fun. The Public LB evaluation on CrowdAnalytix is not absolutely ideal, since you can tune your model to overfit the LB. I still love Kaggle's method of evaluating winners.

Thanks to my family, friends and colleagues (especially my flat-mate and colleague Shashwat) for their help and support, this is a big achievement for me and I'm hoping to perform better in the years to come, and hopefully call myself one of the best Data Scientists of India :-)

Check out My Best CrowdAnalytix Performances

Wednesday, November 12, 2014

Tradeshift Text Classification


The Tradeshift Text Classification competition was held on Kaggle from 2nd October, 2014 to 10th November, 2014.

Objective
Tradeshift has a dataset of thousands of documents, and groups of words are assigned certain labels, eg: Date, Address, Name, etc. The challenge was to create an automated model that predicts which label a certain group of words belongs to.

Data
The train data consisted of 1.7 million rows along with its correct label(s). There were 145 variables containing various attributes about the group of words and also regarding some of the surrounding words/text. There were 33 different possible labels.
The test data consisted of ~ 0.5 million rows for which we had to predict the labels.

Approach/Model
I started off with an absolutely beautiful and elegant online logistic regression code by tinrtgu. Well, so wonderful was this model that most of the top competitors started off and finally used it as part of their best models. Here's the kicker: at the time of sharing the code, it was powerful enough to get into 1st place!

M1
The online model one-hot encodes all the variables, so I rounded off all the numeric variables to one decimal since 'similar' values should ideally be treated the same.

M2
Built Random Forest for y33, the most important label. I took some of the most important variables and tried some interactions with the hash variables for the online code.
Then built RF for all labels individually. Too intensive! Took me over a day across two PCs!

M3
Added the RF-predictions into the online code, and bang! that was it. This submission got me into the top-10.

M4
Tried some other models, but without much success. XGBoost gave promising results, and I added the XGB-predictions into the online code.

Final Model
My final model was M1-M2-M3-M4 into the online code which got me a score of 0.0049356 / 0.0050160 on the LB and would've been ranked in the twenties.

'would've been'? That's right. Abhishek Thakur teamed up with me and we tried some ideas and models together. I don't remember the last ditch ideas that Abhishek tried (I'll update soon), but the final ensemble with a score of 0.0048200 / 0.0048783 certainly helped us secure 13th rank out of the 375 teams.

We tried some other models like GBM, SGD, etc. and some other features, variables, tweaking/tuning of parameters, without much success.

Views
This was the first competition where I came up with a very competitive result and its given me the confidence of coming up with more in the future.

The CV results and LB scores were very close, consistent, and it was an absolutely perfect data-set. The online code is what I'm taking away from this competition, it is now one of my favourite models :-) (Thank You tinrtgu).

Congrats to the Chinese team of rcarson and Xueer Chen and also to the second place team of three French guys who all did a fabulous job of entertaining us to the last day when they had the exact same score on the Public LB! I mean, seriously, how close can you get. Its unfortunate this competition has only one prize, but hey, that's life.

Now my overall Kaggle rank is 206th! Yay! My career-best. You can View My Kaggle Profile.
I'm hoping to get into the top-100 early next year, and hopefully the top-50 (or top-25?) by the end of 2015.

So, lots more to come!

Check out My Best Kaggle Performances

Monday, October 13, 2014

Sudoku Mahabharat 2015


During my trip to London, UK for the World Sudoku Championship 2014, the Indian team had a very entertaining dinner just before the WSC finale and prize distribution. When you sit at a table with Deb, Jaipal, Jayant, Prasanna, Rishi and Swaroop, there's bound to be some fun. Kunal was lost somewhere and Sumit had left to meet some relatives in London. Well, they certainly missed a lot of fun.

During the humorous discussions, we touched upon the topic of creating some more enthusiasm among sudoku solvers in India. It is certainly true that Prasanna, Rishi and me have been the top players in India and people tend to lose motivation knowing its hard to beat us and make it to the team. This was the sixth consecutive year I've been in the India A-Team, and though the top players have been performing extremely well at the World Championships, there was a need to do something more for the players who have potential of becoming the top solvers of India some day. And thus, Sudoku Mahabharat was born.

The idea suddenly hit Deb that we should conduct some sort of national event in India, but more to encourage new players and potential players to participate, win and get a feel of being and performing at the top-level. The idea was discussed and we came up with this event.

Read more about the rules, eligibility, schedule and format of Sudoku Mahabharat 2014-2015

Six of the top players: Jaipal, Prasanna, Rishi, Sumit, Swaroop and me won't be participating, in fact, we'll be organizing the event. This would give some of the upcoming solvers a chance to win this event and hopefully be a part of the Indian team soon.

Episode 1: Standard Variants by Rishi Puri
Dates: 20th-22nd September, 2014
Download Instruction Booklet
Download Puzzle Booklet
View Forum
View Results

Episode 2: Irregular Variants by Rohan Rao
Dates: 18th-20th October, 2014
Download Instruction Booklet
Download Puzzle Booklet
View Forum
View Results

Episode 3: Odd-Even Variants by Deb Mohanty
Dates: 15th-17th November, 2014
Download Instruction Booklet
Download Puzzle Booklet
View Forum
View Results

Episode 4: Outside Variants by Rakesh Rai
Dates: 20th-22nd December, 2014
Download Instruction Booklet
Download Puzzle Booklet
View Forum
View Results

Episode 5: Math Variants by Prasanna Seshadri
Dates: 17th-19th January, 2015
Download Instruction Booklet
Download Puzzle Booklet
View Forum
View Results

Episode 6: Neighbours by Rajesh Kumar
Dates: 21st -23rd February, 2015
Download Instruction Booklet
Download Puzzle Booklet
View Forum
View Results

Episode 7: Math Variants by Swaroop Guggilam
Dates: 21st-23rd March, 2015
Download Instruction Booklet
Download Puzzle Booklet
View Forum
View Results


Whether you're a beginner, regular or pro, whether you're 17yrs, 31yrs or 67yrs, whether you're an Indian or not, this is a chance for you to participate (even for fun!) and solve some different interesting sudokus created by some of the best sudoku creators in India. There is something for everyone to enjoy! Who knows, maybe you can be the one who wins the Sudoku Mahabharat !

Sunday, August 24, 2014

World Sudoku Championship 2014

The 9th World Sudoku Championship was held in London, UK in August, 2014.

Official Website

The Indian Team for the WSC were the winners of the Times Sudoku Championship: Prasanna Seshadri, Rishi Puri, Sumit Bothra and me.


Prasanna, Rishi and me were in the team last year as well, and we secured 15th, 28th and 16th rank respectively, so, we were hoping to do better this year, in the individuals as well as the team (India were 7th last year). We had a B-Team too, Swaroop, Jaipal, Jayant and Kunal.

I've been on a long break from solving puzzles, but after ISC, I decided to prepare for the WSC. I received a jolt with a pathetic TSC performance, but I was glad I made mistakes in TSC, rather than ISC and hopefully WSC.

When the WSC IB was released, I was a little disappointed. And after the WPC IB was released, I spent all my time in preparing for WPC. Well, so much that I organized the WPC (World Practice Championship) on LMI.

So, what was wrong with WSC? Nothing wrong, it was just that the WSC looked a little bland. All the usual variants, some repetition among types, no hard Classics and nothing specifically to practice or prepare for. The rounds felt easy-ish and finish-able, and it did turn out that way.

Round 1 was good, I missed a couple of low-pointer Classics.
Round 2 was bad. I messed up the Surplus, missed the Killer (and later had a terrible error in Diagonal).
Round 3 was good, just missed the Max Triplet.
Round 4 was excellent, just missed the Inequality (and later had a single cell error in a Classic).
Round 5 was ok, got stuck at Killer Pro and lost some time, but managed a decent score.
Round 6 was a round I was dreading, and I didn't do well. Wasted time on the Toroidal that went wrong towards the end. Few more seconds and I could've finished the Parquet.

Round 7 was a team round, where we finished with bonus. We made a single digit error and lost half the puzzle points and all the bonus :-|
Round 8 was a team round that we royally screwed up. Well, we certainly didn't prepare enough as a team.

So, that was Day-1. I was 12th after the Day-1 results were out. Felt OK-ish. Playoffs chances were bright, but had to do well on the second day.
Prasanna had a terrible Round 5 and Rishi had an average day, pushing them outside the top-20.

I was all set for Day-2.

Round 9 was the 'big round', the longest round with maximum points. During last year's WSC, there was a similar 'big round' on the morning of the second day, which I totally messed up and dropped a few ranks. This year, I didn't mess up, but I didn't do my best either. I broke Clone and Cylindrical. Didn't have time for Between. And later had a mistake in Diagonal. Why can't I solve Diagonal Sudokus this year?
After this round I knew I didn't have any chance for playoffs. So, my goal was to at least improve on last year's rank (16th).

Round 10 was the last round, an overlapping sudoku (similar to last round of WSC 2012), and I like these. I finished WSC on a high by completing this round in 8mins, securing a 12-min bonus. Only 3 players finished better, with a 13-min bonus.

Round 11 was a team round where we had to solve sudokus that were linked. We were confident of doing well in this round since we all had 'different' variants that we liked, and we did well.

So, that was my WSC 2014. I stood 14th. Better than last year, but still a big gap between 10th and 14th. I'm happy I've been performing consistently over the years, with/without practice: My last 5 WSC ranks are: 14, 16, 8, 12, 15

Prasanna never really recovered from that bad round 5, and finished 21st. Rishi's performance was below par, 36th. Sumit made a couple of mistakes towards the end, and dropped out of the top-50, with his 55th.

Our Team was swapping between 7th and 8th over the rounds, and we thought we'll be 7th (again!). But, special thanks to the French team for making a goof-up in the last team round. They lost nearly 2000 points, and that pushed us to 6th, which is India's best team rank ever.



The playoffs format was different from last few years, and even though there were 10 in the playoffs, it was mainly a fight between the top-4. I found this format better than most other WSCs especially since, the worst case scenario for the preliminary toppers is 4th place.

Tiit and Kota were going head-to-head, it was a very close contest between the two, which Kota eventually won. (If he had lost, he would be WSC runner-up for the 4th consecutive year!). On the other hand, Tiit has won the preliminary twice but still can't call himself a WSC Champion. Well... that's life. Maybe next year.
(There was some problem with Bastien's and Jakub's papers during the playoffs, but I think, the organizers found a way out by awarding them joint 3rd place)

Congrats to Kota Morinishi for winning WSC 2014, Tiit Vunk for 2nd and Bastien-Vial-Jaime and Jakub Ondrousek for joint 3rd. They have been performing well consistently and they are the deserving winners.

View the complete WSC Results

Overall, I think this was a much easier WSC than the last few years. I enjoyed the sudokus, the rounds and the experience. The team rounds were fun and easy. The playoffs were good and fair. So, that brings an end to a good and successful WSC, but a little bland one. Yes, the only little negative is that too many standard variants, and some unnecessary repetition. But, I guess that's what made it feel easy and fun :-)

Thanks to UKPA and the UK Organizers and Volunteers for conducting this wonderful WSC. Thanks to all the authors for the enjoyable sudokus and hope to see more in future.

This was my 6th WSC, and it would be in my favourite two.

Monday, July 21, 2014

Beginners Contest on LMI



I authored the July Beginners Contest on LMI. This was a short contest consisting of 4 Classic Sudokus and 4 Sudoku Variants. The difficulty was aimed at ensuring new comers to be able to solve these. The contest was held from 23rd July - 28th July, 2014.

The Sudoku Variants that appeared are Arrow, Diagonal, Extra Region and Trio. You can view some examples here.

You can discuss about the contest on the LMI Forum.

There were totally 295 participants across 38 countries! Congrats to sworls (USA), Mehmet Eren (Turkey) and Vidhya (India) for being the top-3 Beginners. And congrats to Kota Morinishi (Japan), Timothy Doyle (France) and Hideaki Jo (Japan) for being the top-3 Seasoners.

You can view the Complete Results.

Overall from the feedback I received, most participants enjoyed the 'easy' Classics. Among the variants, Trio is always easy, Diagonal was easy, Extra Region was medium and Arrow, well, Arrow was something, right? The Arrow Sudoku was certainly challenging, especially for beginners, but I don't think it was too tough :-)
Maybe you can decide for yourself after checking the solve below:

This is the puzzle:


The 'long' arrow has to be '6'. After that, I admit, the next step is not straightforward. Intuitively, I thought people will target the bottom-left, and there is an opening there. You should get the following pencilmarks using simple addition rules and constraints.


If you look at R8C4, it cannot be '2' or '3' because it will result in the following contradictions for those two arrows. This is not very easy to identify, but filling those two arrows is quite tight.


Hence, R8C4 is a '1'. From here, the solve is smooth.



At this stage, R8C3 cannot be '3', hence it is '2' and that will also enable the other arrow to be completed. There are other ways too, but if you try filling up box 7, along with the arrows, there is just one possibility.


Using Classic rules, you can reach this stage, with those two possibilities for '5' in box-8.


The arrow in box-2 covers R1C5 and R2C6. Minimum in R1C5 is '3' and minimum in R2C6 is '6'. So, it has to be 3+6=9, which will also give you the 5-9 pair in box-8.


The centre box can then be completed, which will give you a few more digits using Classic rules.


The arrow with '7' can only be 1-6 and you get a few more digits.


You have just one arrow left to complete, and eliminating '4' is simple since neither 1+3 nor 2+2 is possible. Using '8', you should get 3+5 and the rest gets solved using simple Classic rules.


So, what do you think now? Tough? Maybe not that much!

This Arrow Sudoku was not very trivial and a lot of solvers struggled on this variant. But I hope you enjoyed solving it at the end.

And, I hope everyone enjoyed the sudokus in general, and thanks for participating in this contest! :-)

Monday, July 14, 2014

Happy Birthday Nanamma!

I created an Alphabet Sudoku for my Grandmother's birthday themed on her name. Her name is Shanta Murthy and she loves solving all kinds of puzzles. She solves most of the online puzzle contests at leisure during her spare time and usually completes solving all puzzles, sometimes taking multiple weeks for difficult ones! (I myself give up on really tough puzzles :-) )

She turned 82-yrs on 14th July, 2014 but has a puzzling mind of a 28-year-old. She feels proud of all my puzzle-related achievements but is too shy to compete herself. I keep telling her she will easily win an over-75-years category prize (if there was any!) and I still hope some day she does.

I lovingly call her 'Nanamma' (which means 'paternal grandmother' in our local language) and I dedicate this Alphabet Sudoku to her.



Nanamma solves more puzzles than me every week and it has helped keep her mind active and agile through these years. I hope she continues to enjoy solving puzzles and maybe some day she'll create one for me :-)