Search This Blog

Friday, June 24, 2016

Indian Puzzle Championship 2016


The Indian Puzzle Championship (IPC) 2016 was held on 17th July, 2016 in Chennai.

The championship was an offline event (after four years of having it online), where the top-50 qualifiers of Puzzle Ramayan were invited. Prasanna is undoubtedly the best puzzle solver in India today, and has been last 3yrs.

View Top Qualifiers

Last 2yrs were exciting competing with Amit. He won the 2014 edition and I came in second, while I won the 2015 edition and he came in second. I was expecting him to be the biggest challenger for the title, since Prasanna, being the best Indian at the World Puzzle Championship last year, gets a wild card for the national team this year and hence, decided to organize the national event.


Puzzles
The championship consisted of four rounds. When you have Deb and Prasanna authoring puzzles, you know you're in for a treat. Really nice set of puzzles.

The 3rd round Sprint was my favourite, and also, the only round I topped.
The 4th round was Casual-type puzzles, many of them visual, and supposedly one of my strengths. But, I did horribly.

Those of you who'd like to solve these puzzles can purchase it here (Unfortunately due to certain reasons, they are not publicly available this year).


Results
I expected it to be a close fight between Amit and me. But, after the first two rounds, Amit had a big enough lead over me and it was pretty hard to catch up from then. I did cover up a few points in the third round, but had a terrible fourth round.

Amit won his second IPC and I stood 2nd (for the third time, after 2009 and 2014).

The ever-consistent Rakesh stood 3rd with a strong performance. Ashish's performance was disappointing, but just made it into the B-Team with his 6th place finish.

View Complete Results

The winning trophies were really cool! Custom-made puzzles by Prasanna (and one by me, the second one, which I ended up getting :-| ) printed on the trophy with the themes '1', '2' and '3'. Thanks to Sumit Bothra for getting these made.




Organization
Thanks to Deb Mohanty and Prasanna Seshadri for organizing this event smoothly. Also, a big thanks to Varun, Ezhilarasi, Ashish, Kumaresan, Rakesh, Kishore, etc. and the other folks in Chennai who helped out with the event. It was a big success!

Indian Sudoku Championship 2016


The Indian Sudoku Championship (ISC) 2016 was held on 16th July, 2016 in Chennai.

The championship was an offline event (after three years of having it online), where the top-50 qualifiers of Sudoku Mahabharat were invited. This is the first time the reigning champion did not defend the title. Rishi Puri, who won ISC in 2014 and 2015, has called it quits for his sudoku career.

View Top Qualifiers

Last 3yrs were very exciting competing with Prasanna and Rishi. Its unfortunate that neither of them participated this year since Prasanna, being the best Indian at the World Championships last year, gets a wild card for the national team this year and hence, decided to organize the national event.

I've traditionally done well in offline ISCs... in fact, only 3 of the ISCs were offline (2010, 2011, 2012), and those were the 3 times I won! This year I made it 4/4.


Puzzles
The championship consisted of four rounds. When you have Deb and Prasanna authoring sudokus, you know you're in for a treat. Really nice set of puzzles.

The fourth round of 6x6 sudokus had a fun twist to it, and was well thought of. The timings were perfectly set and the event ran very smoothly.

Those of you who'd like to solve these puzzles can purchase it here (Unfortunately due to certain reasons, they are not publicly available this year).


Results
I'm glad I topped all four rounds and won my fourth ISC title, this one after 4 years!

Rakesh Rai pipped Kishore Kumar for 2nd place as Kishore had a really bad 4th round. Most of the other results were as expected. Akash Doulani and Pranav Kamesh were the top inexperienced players comfortably, and I'm glad they'll finally be making it for their first World Championships.

View Complete Results

The winning trophies were really cool! Custom-made sudokus by Prasanna printed on the trophy with the shapes '1', '2' and '3'. Thanks to Sumit Bothra for getting these made.






Organization
Thanks to Deb Mohanty and Prasanna Seshadri for organizing this event smoothly. Also, a big thanks to Varun, Ezhilarasi, Ashish, Kumaresan, Rakesh, Kishore, etc. and the other folks in Chennai who helped out with the event. It was a big success!

Thursday, May 5, 2016

The Seer's Accuracy


AnalyticsVidhya organized a weekend hackathon The Seer's Accuracy on 29th April - 1st May, 2016.

In the midst of a new job, new city, I wasn't sure if I'll get enough time to participate in this hackathon. But fortunately, it was a relatively light weekend.

Problem
The challenge was to predict which customers would be return customers to a chain of stores. Looking at it another way, it was predicting which customers would churn.

Data
The train data consisted of customers (a.k.a. clients) and their transaction history in the years 2003 - 2006. The test evaluation was on which clients would return in 2007.

There was no test data per se, and it turned out to be the most crucial part of this challenge.

Overall, very clean data and very interesting problem. Kudos to AV!

Model
Right from the beginning I felt setting up a validation framework is going to be the key. And with a few LB submissions, I realized it was going to be extremely important to have a good validation set too.

I started off just like most other participants by using 2003-04-05 as build set and 2006 as the validation set and running a CV on it.
What finally catapulted me up the LB was when I added 2003-04 as build and 2005 as the validation set as well in my CV framework.

I think this resulted in a much stabler validation set and my CV and LB improvement was much more in sync.

Since the variables were limited, I treated and tested each of them individually and finally had a model with 335 features.

My final model was a blend of 3 XGBs on varying subsets of data and features.
It was a very minor improvement over my single best model.

GitHub
View My GitHub Repository

Results
I stood 1st on the public LB scoring 0.8856 and 1st on the private LB too, scoring 0.8800 using the AUC metric with the username 'vopani'.

Congrats to orenov/DataGeek for 2nd place and Bishwarup for 3rd place.

View Final Results

Views
Feels good. Really good.
Not just for winning, but for building a solid architecture which enabled a strong and stable model resulting in a considerable lead over the rest.

And this is also my first win on AV! :-)

Thanks to the AV organizers for this hackathon, was top quality and totally worth spending a weekend over.

View AV article on winners.

Sunday, March 6, 2016

Telstra Network Disruptions


The Telstra Network Disruptions competition was held on Kaggle in Nov, 2015 - Feb, 2016.

Objective
The objective was to predict the severity of a service disruption (whether it is a momentary glitch or a serious interruption of connectivity) on the Telstra network.

Data
The data consisted of disruptions along with features related to logs, events, resources, severity types across various locations.
The target variable was the severity of the disruption, into 3 classes.

Model
There was a golden insight in the data, and exploiting that became a very interesting challenge.

I ensembled several XGBoosts, on different subsets of the data and features, and some combinations of parameters.

The features used were the one-hot encoded raw features along with some interesting features built using the golden insight.

GitHub
View My GitHub Repository

Results
I stood 10th on the public LB and 9th on the private LB, scoring 0.40735 / 0.40267 using the logloss metric. My username is 'Vopani' and I competed as 'Anonymous Ghost' during the competition.

Views
This contest was all about finding and hacking that golden feature. The feature was nothing complex, it was simply the ordering of the observations that mattered. It was hard to spot because the ordering held true in the feature files (log, event, resource, etc.) and not on the original train/test data.

After identifying the relevance of ordering, it was very interesting to work on feature engineering and build features to improve the model without overfitting.

I'm glad I finished in the Top-10 and it also becomes my first Top-10 finish on Kaggle as an individual. I've moved to 70th in overall Kaggle rankings. I should easily be able to get into Top-50 by the end of the year. Maybe even Top-25.

Check out My Best Kaggle Performances

Saturday, February 13, 2016

AirBnB New User Bookings


The AirBnB New User Bookings competition was held on Kaggle in Nov-15 to Feb-16.

Objective
The objective was to predict in which country a new user on AirBnB would make their first booking.

There were 11 potential countries along with a 12th class - NDF (No Destination Found), indicating the user did not make any booking.

Data
The data consisted of user characteristics like language, age, browser, date-of-account-creation, OS, etc. for the train and test users.

There was data on the actions taken by users on the website along with the details of the action and duration.

Model
It was evident that the best way to quickly get a good score was to focus on classifying the NDF vs non-NDF users. So, I built a Logistic Regression on the one-hot encoded action features from the sessions data as a binary classifier for NDF vs non-NDF, only considering users present in the sessions data. This was the base classifier.

I then built a meta classifier using, well, everyone's favourite nowadays, XGBoost. It used the raw user features, along with the one-hot encoded features from sessions data, and finally, the LR predictions.

I did not complicate the model or ensemble too much due to lack of time, and also since the CV and LB were not perfectly correlating. Hence, I chose fairly simple models with some feature engineering.

GitHub
View GitHub Repository for the complete code, results and output.

Results
This model scored 0.88081 on the public LB which was ranked 89 and scored 0.88625 on the private LB which was ranked 23.
The metric used was NDCG.

View Public LB
View Final Results

Views
It was a very interesting dataset, and a good practise in building features from the sessions data, and without that, it wasn't possible to get a good score. It was disappointing that I had so many ideas which involved a lot more time to try out and code, but wasn't able to.

So, I think it was a simple stable model with lesser overfitting compared to many other competitors who dropped on the private LB.
In the end, I'm happy with the result, and this improves my overall Kaggle rank to 96th. So, finally I get into the Top-100 and on the first page of the rankings :-)

Hoping to improve on this further this year, and hopefully get into the Top-50 or Top-25 some day.

Check out My Best Kaggle Performances

Thursday, February 11, 2016

Puzzle Grand Prix 2016


The WPF Puzzle Grand Prix 2016 is here! After successful editions in 2014 and 2015, the 2016 edition consists of eight rounds held across the first seven months.
The top-10 finalists will be invited for the GP playoffs during WPC 2016 (Slovakia).

The format has changed a bit this year, with the contest having two sections. A Competitive section which is the main section on which the toppers will be decided, and a Casual section, comprising of more 'culture-neutral', non-grid based puzzles, geared towards leisure solvers. Its very unlikely competitors will be able to participate in both sections, so, I guess most players need to choose one.

I think I'm better at the 'Casual' sort of puzzles, and hence, will be competing only in that section.

Scoring System
There has been some discussions regarding the new scoring system for the GPs. Historically, normalization of scores for a championship consisting of various rounds has worked well due to the unreliability of having similar rounds in terms of scores and difficulty.
The GPs did use normalization in the previous editions, and it was universally accepted.

I'm not convinced there was a need to go ahead without normalization this year, so it remains to see whether or not it will work. You can find some pros/cons being discussed about the new scoring system on the GP Forum.
Being part of the organizing team of Sudoku Mahabharat / Puzzle Ramayan (which are very similar to the structure of GPs), we had to change the scoring system during this year's rounds, due to the inconsistencies without normalization. So, I'm not particularly in favour of dealing with raw scores.

View Championship Page
View Current Rankings



Round 6: Serbia (10th - 13th Jun, 2016)
As the competition is getting heated up, I was fairly comfortable with the puzzles in this round. Enjoyed the set and managed to finish all but one puzzle (Weights).

Before this round, I was leading with ~ 25points over Adam Bissett and ~ 50 points over Yuhei Kusui. Ironically, all three of us scored the exact same points: 344, in this round, keeping the standings the same.

It now boils down to the last two rounds to decide the winner.


Round 5: USA (13th - 16th May, 2016)
I was afraid that I barely managed to score 200 points in this hugely big 500+ pointer round. But it seems like it was too hard and everyone struggled.

The concept of Escape The Grand Prix is really nice, but not suitable for a time-contrained online puzzle round. This is easily going to be the discarded round for most of the top players.

Kudos to Randy Rogers for scoring 254 points, and Sinchai Rungsangrattanakul for scoring 242, way above the rest of the lot, while I scored just 205. Which is my poorest round so far.


Round 4: Hungary (15th - 18th Apr, 2016)
An easy set here. Managed to finish all and score full points, thus keeping my lead intact.
Nice puzzles too.

The four rounds so far have been worth 293, 402, 420 and 259 points. What in the world is the logic of not having normalization. That is the biggest failure of this year's GPs.


Round 3: Germany (18th - 21st Mar, 2016)
What an amazing set of puzzles. Just wow! This is the best set of puzzles I've solved in a long time... each and every one of them is a beauty.

The Instructionless Machine puzzles... exceptionally interesting and well-made. I think it was the sheer fun of the round that made me perform so well. I topped with 420 points, scoring over 100 points more than everyone else, except Jarett Prouse who scored 382.
Which means, I'm back in the top-3.


Round 2: Slovakia (19th - 22nd Feb, 2016)
Bad round. Lost time on the scrabble puzzles and wasn't able to complete them at the end.

Scored 283 points which is bad compared to the highest being 402. Hopefully, this could be one of my discards.

With no normalization, we have two rounds, one worth 293 points and the other worth 402 points. I just don't get it.


Round 1: India (22nd - 25th Jan, 2016)
It was a spur of the moment decision to participate in this round, since, being authored by Indians, I just assumed I couldn't compete. But Prasanna informed me that I was the test-solver only for the Competitive section and I could, in fact, participate in the Casual. And I did.

I finished 4th with 262 points. The highest score was 275 points by Adam Bissett of UK. Not a bad start.

The puzzles were excellent. The Buttons and the Number Series really got me scratching my head for a long time, and ultimately ended up missing out on two Number Series and a minor error in Shape Count.

Its quite funny that considering normalization is not being used this year, you'd expect the rounds to at least be similar in terms of points. The Competitive section was worth 697 points while the Casual section was worth 293 points. So, I have no clue where this is going. Lets hope its not too bad.

Wednesday, January 27, 2016

Sudoku Grand Prix 2016


The WPF Sudoku Grand Prix 2016 is here! After successful editions in 2014 and 2015, the 2016 edition consists of eight rounds held across the first seven months.
The top-10 finalists will be invited for the GP playoffs during WSC 2016 (Slovakia).

Due to certain unfortunate incidents, I wasn't able to compete completely in the first two editions of the GP. Well, I'm really hoping to make it this year.
Rishi Puri, the current and two-time national champion, featured in the playoffs in both years while Prasanna Seshadri was in the playoffs last year. With Rishi 'retiring' from active participation, it'll be interesting to see how the GP unfolds for the Indians!

Scoring System
There has been some discussions regarding the new scoring system for the GPs. Historically, normalization of scores for a championship consisting of various rounds has worked well due to the unreliability of having similar rounds in terms of scores and difficulty.
The GPs did use normalization in the previous editions, and it was universally accepted.

I'm not convinced there was a need to go ahead without normalization this year, so it remains to see whether or not it will work. You can find some pros/cons being discussed about the new scoring system on the GP Forum.
Being part of the organizing team of Sudoku Mahabharat / Puzzle Ramayan (which are very similar to the structure of GPs), we had to change the scoring system during this year's rounds, due to the inconsistencies without normalization. So, I'm not particularly in favour of dealing with raw scores.

View Championship Page


Round 3: Czech Republic (4th - 7th Mar, 2016)
Coming soon...


Round 2: Serbia (5th - 8th Feb, 2016)
Ohhh no. I bombed this round. Just had a bad day. And another submission mistake made it worse. So, that makes it two bad rounds out of two :-(

Puzzles were nice, nothing extraordinary. Many of them turned out to be Converse-like variants, which incidently comes right before the Converse round of SM, so good practise there.


Round 1: Netherlands (8th - 11th Jan, 2016)
Not the start I was hoping for. Had a rough solve, got stuck up time and again. To make things worse, I had one submission incorrect.

Tiit finished the set in just less than an hour, which is phenomenal. Prasanna finished the set in 79mins putting him in 5th. I finished in 84mins, but the mistake dropped me to 14th place.
The puzzle quality was excellent, as expected from the Dutch authors.

I hope this is the round that gets discarded!

Monday, January 4, 2016

Classic Tapa Contest 2016


Classic Tapa Contest (CTC) 2016 was held in Jan-Feb, 2016 on LMI.

Championship Page

View Forum
View Results

1. sai (Japan)
2. EKBM (Japan)
3. Prasanna16391 (India)
4. deu (Japan)
5. Psyho (Poland)
6. Para (Netherlands)
7. nyoroppyi (Japan)
8. willwc (USA)
9. kiwijam (New Zealand)
10. anderson (USA)

View Complete Results

sai won the last CTC and this time completely dominated at the top. After a horrific start, Endo blew through the middle days and just managed to overtake Prasanna in the last week to finish 2nd. And phenomenal performance by Prasanna, who takes 3rd, after an extremely consistent run of two months.

I was in the race to get into the Top-20, but unfortunately missed a few days due to personal matters. Nevertheless, I'm quite happy with my performance and enjoyed the Tapas.

I'm sure a lot of people are going to miss CTC... it just gets on you after 50 days :-)

Wednesday, December 16, 2015

Rossmann Store Sales


The Rossmann Store Sales competition was held on Kaggle in Nov-Dec, 2015.

Objective
Rossmann operates over 3000 drug stores in 7 European countries. Currently, Rossmann store managers are tasked with predicting their daily sales for up to six weeks in advance. Store sales are influenced by many factors, including promotions, competition, school and state holidays, seasonality, and locality. With thousands of individual managers predicting sales based on their unique circumstances, the accuracy of results can be quite varied.

The objective was to predict the sales of various Rossmann stores in Germany.

Data
Train data consisted of sales from over 1000 Rossmann stores along with information related to promotions, competitions, holidays, etc. upto July, 2015.

Test data consisted of dates in August and September, 2015 for which we had to predict the sales.

Approach
Being a classic time series sales forecasting problem, I explored two approaches. One being the standard building of tree-based and linear models. The other being trying out time series models like ARIMA.

It became quickly evident from cross-validation and validation results that ARIMA wasn't working. XGBoost was giving much better results.

There was a lot of external data shared and available, but none of those made a big improvement in the model. My final model didn't use any external data either.

Building models at a store-level was not giving as good results as building a model using all the data together, but it helped while blending models.

Model
I built multiple XGBoost models on different subsets of the entire data and averaged them. I merged these with store-level models of XGBoost, Random Forest and GBM. The blending of models gave a huge improvement and ultimately lead to the stability of the predictions.

I finally tweaked the predictions by using a multiplicative factor of 0.98 to get the best fit to the LB.

I usually share my code on GitHub, but this time I decided against it, since I haven't done anything extraordinary or special.

Results
My model gave a RMSPE of just below 0.10 on the public LB with rank 66th and RMSPE of just below 0.11 (in fact, I scored 0.10999!) which ranked me 14th on the private LB out of 3303 teams.

A lucky jump, having chosen a stable model, which results in my best individual performance on Kaggle till date, improving on my 14th rank / 2256 teams in the TFI competition.

View Public LB
View Final Results

Views
It was a tricky contest, mainly due to the nature of the public and private LB split. It was overwhelming to see so much external data being shared and used. Maybe under other circumstances, this could have played a much more important role.

Congratulations to the winner, Gert, who performed fantastically, by being way ahead of the lot in the public LB with very few submissions! And finally being stable enough to win on the private LB, again with a big lead.

So, I gained some good points from this contest, and moved to 111th in overall Kaggle rankings. My year-end goal was to be in Top-100. I'm close, and with the Walmart contest left, I might just make it.

Check out My Best Kaggle Performances

Monday, November 23, 2015

Black Friday Data Hack


AnalyticsVidhya organized a weekend hackathon called Black Friday Data Hack, which was held on 20th-22nd November, 2015.

Black Friday is actually the following weekend, but that's when we've to relax and enjoy :-)

The last hackathon was quite disappointing due to the randomness in the data and the evaluation metric. I was hoping this one would be better.
And it was. Much better.

Problem
The challenge was to predict the purchase amount of various products by users across categories given historic data of purchase amounts.

Data
In general, when you have more data, its always better. The train data had ~ 5.5 lakh observations and the test data had ~ 2.3 lakh observations. The data was very very clean and it feels wonderful to work on such datasets.

The data was of users who purchased products with the amounts. The products had data on three types of categories. The users had data about their age, gender, city, occupation, locality and marital status.

We were to build our models on the train data and score the test data which had pairs of user-product not present in the train data. The evaluation metric was RMSE, which also seemed a very appropriate choice for this problem.

Approach
I spent the first few hours just exploring the data, summarizing variables, plotting graphs, playing around with pivots and in parallel building base models (of course, XGBoost).

On the first day, I was able to go below 2500 with an optimized XGBoost model on raw features. It got me into the Top-3 and since then I've managed to maintain a position in the Top-5.

While checking the variable importance of my XGBoost, I found Product_ID was the most important variable and intuitively it made sense. So, I just submitted the average purchase amount of each product and voila! it scored 2682, which didn't seem like a very bad score. So, all those of you who couldn't cross 2682, here's a simple solution you missed.

Usually ensembles win competitions, but since I couldn't get any model close to the performance of XGB, so I decided to challenge myself to build a single powerful model. Which means, feature engineering.
These two days gave me some wonderful insights on how powerful feature engineering is. With some analysis, gut, trying, cross-validating, here are my final set of features that I used:

Model
User_ID: Used as a raw feature

User_Count: Number of observations of the user

Gender: Converted to binary

Age: Converted to numeric

Marital Status: Used as raw feature

Occupation: Used as raw feature

City Category: One-hot encoded features

Stay In Current City: Converted to numeric

Product Category 1, 2, 3: Used as raw feature

Product_Count: Number of observations of the product

Product_Mean: Average purchase amount of product

User_High: Proportion of times the user purchases products at a higher amount than the average purchase amount of the product

I built an XGBoost with these features, and the code is open-sourced on GitHub, the link is given below.

One very interesting feature I built was
F_Prop: Average purchase amount of product by female users / Average purchase amount of product by male users

This was among the top-3 important variables and gave a CV of ~ 2419 but the LB remained very similar ~ 2430, so I wasn't sure about it. I decided to go without this.

GitHub
View GitHub Repository

Results
This model gave me CV score of ~ 2425 and public LB score of 2428. I was 4th on the public LB, with Jeeban, Nalin and Sudalai in the Top-3. And we finished in the same positions with my final rank being 4th in the private LB.

Views
This is one of the best data-sets I've worked on in a while. The CV and LB scores were perfectly in sync and it was very satisfying to build features and improve the CV as well as LB scores. I'm happy with my performance as I managed to squeeze quite a lot of from the data with a single model.

I might have done better with an ensemble, but just couldn't get anything to work well. And after a while, was just too tired.

Overall, a great weekend, mostly spent on my laptop. For those of you who had memory issues, I worked on my 4GB MacBook Air throughout the weekend. Algorithms and models will advance and become optimized every day, but the power of building good features is still in the hands of Data Scientists like us.
Make the most of it until the machines come and take over ;-)

Thanks to all the folks at AnalyticsVidhya for organizing this hackathon. A big thumbs up from me.

Looking forward to the next Hackathon, and hope it gets better and more competitive.

External Links
View Other Players' Approaches on AnalyticsVidhya
View 3rd place solution code on GitHub by Sudalai Raj Kumar
View 5th place solution code on GitHub by Aayush Agrawal