Search This Blog

Wednesday, August 3, 2016

The Smart Recruit


AnalyticsVidhya organised a weekend hackathon called The Smart Recruit, which was held on 23rd-24th July, 2016.

I won the previous hackathon, The Seer's Accuracy, and was hoping to do well in this one too.

Problem
The problem was to identify which agents would be successful in making sales of financial products. So, it was a binary classification problem.

Data
Like the previous hackathon, the data seemed quite good and promising.

Train and test data consisted of agent applications with data about the application, the manager and a few related features about them.

Model
I'm sure most participants just went ahead and dumped the data into XGB-type models with a lot of scores hovering in the 0.63 - 0.66 AUC range.

I tried to get a robust/stable validation framework, like I mentioned in AV's article on Winning Tips. Didn't seem to work/help. The CV/LB scores were all over the place in my first few submissions.

Thats when I decided to take a step back and inspect the data in detail. It was evident to me that there would be a huge LB shake-up due to the variance between the CV-LB scores. Hence, didn't make much sense to spend too much time on the data trying to optimize models. Instead, I tried to look for some pattern/feature which could boost me score over the expected error margins.

And that's exactly what happened. A simple plot of the target variable showed a pattern, which seemed too good to be true. I tried a feature using this and my CV jumped to 0.8... and that was the feature that ultimately proved to be the winning one.

Here's the plot that changed everything:

This is the plot of the target variable for the first four days. A clear pattern exists where you see most of the 1's at the beginning of a day and most of the 0's at the end of the day. You can plot the target variable of any single day and observe a similar trend.

Leakage? Possible. Hidden trend? Possible. At first I was convinced it was leakage and a data preparation issue, but later, felt there was a possibility that applications received towards the end of the day are more likely to be rejected than ones received early.

Either ways, I polished this feature using Order_Percentile in my code, which was the most important feature.

My final model was a single XGBoost with 14 features, with the other 13 being cleaned up features from the raw variables. I achieved a CV of 0.887 which was in the same range as the LB. I'd have liked to try out some more parameter tuning and ensembling, but with the limited duration of a hackathon, there wasn't any time left.

GitHub
View My Complete Solution

Results
I stood 1st on the public LB with 0.885, with good friend and rival competitor SRK in 2nd, who teamed up with Kaggler Mark Landry, with 0.876 and another team of Kanishk Agarwal and Yaasna Dua in 3rd with 0.839. No other team figured out the winning feature and their scores were below 0.71.

The rankings held same on the private LB, but it was much closer, with SRK-Mark scoring 0.7647 and I scoring 0.7658.

My username is 'vopani'.

View Complete Results

Views
My 2nd AV win on the trot and while not the best way to win it, I'm happy I could find a useful winning feature in the data.

Congrats to the ever consistent SRK, who also happens to be someone I'm chasing on Kaggle :-)

Fun weekend, bonus to win it, and looking forward to the next hackathon, where I'll be on a hat-trick!

An interesting co-incidence: I got the exact same score on the public LB (0.8856) in the previous hackathon too, The Seer's Accuracy !!!

External Links
View AV article on the winners
View 2nd place solution by SRK
View 3rd place solution by Kanishk Agarwal

Friday, July 29, 2016

Nibbl


Nibbl is a platform for authoring and solving various grid-based logical puzzles. Being a part of the puzzling world, having authored, organized and participated in various international puzzle championships across the globe, I've wondered what is the best way to reach out, publicise and improvise on these popular genres across different channels and audiences. With the advancements of mobile technology, it is one of the fastest and most convenient forms of information distribution.

So, why Nibbl?

Yes, there are a few puzzle solving (especially sudoku) apps out there, but most of them are computer generated puzzles built by programmers. Don't we all love hand-crafted theme-based puzzles with wonderful solving paths to 'tickle our minds'? All puzzles on Nibbl are custom created by authors.

One of the best features in Nibbl is the ability to view the author's intended solving path. A great way to learn from the creators on what logic was applied when and how.

Currently, it has a few popular puzzle genres but more will be added soon.
Nibbl Authors can publish puzzles as per their choice and liking and decide their own pricing. Creating puzzles is a matter of few taps and can be done on the app itself. Interested authors can contact nibble.appfactory@gmail.com.

Solvers can download the app from:
Google Play Store: https://play.google.com/store/apps/details?id=rao.rajnikant.ips
Apple Store: https://itunes.apple.com/in/app/nibbl-solvr/id1132803199

Feel free to use my referral code for some bonus credits: ROHA3286

I plan to publish some puzzle packs on Nibbl later this year.

P.S. My father, Rajnikant Rao, is the main driver of the project :-)

Friday, June 24, 2016

Indian Puzzle Championship 2016


The Indian Puzzle Championship (IPC) 2016 was held on 17th July, 2016 in Chennai.

The championship was an offline event (after four years of having it online), where the top-50 qualifiers of Puzzle Ramayan were invited. Prasanna is undoubtedly the best puzzle solver in India today, and has been last 3yrs.

View Top Qualifiers

Last 2yrs were exciting competing with Amit. He won the 2014 edition and I came in second, while I won the 2015 edition and he came in second. I was expecting him to be the biggest challenger for the title, since Prasanna, being the best Indian at the World Puzzle Championship last year, gets a wild card for the national team this year and hence, decided to organize the national event.


Puzzles
The championship consisted of four rounds. When you have Deb and Prasanna authoring puzzles, you know you're in for a treat. Really nice set of puzzles.

The 3rd round Sprint was my favourite, and also, the only round I topped.
The 4th round was Casual-type puzzles, many of them visual, and supposedly one of my strengths. But, I did horribly.

Those of you who'd like to solve these puzzles can purchase it here (Unfortunately due to certain reasons, they are not publicly available this year).


Results
I expected it to be a close fight between Amit and me. But, after the first two rounds, Amit had a big enough lead over me and it was pretty hard to catch up from then. I did cover up a few points in the third round, but had a terrible fourth round.

Amit won his second IPC and I stood 2nd (for the third time, after 2009 and 2014).

The ever-consistent Rakesh stood 3rd with a strong performance. Ashish's performance was disappointing, but just made it into the B-Team with his 6th place finish.

View Complete Results

The winning trophies were really cool! Custom-made puzzles by Prasanna (and one by me, the second one, which I ended up getting :-| ) printed on the trophy with the themes '1', '2' and '3'. Thanks to Sumit Bothra for getting these made.




Organization
Thanks to Deb Mohanty and Prasanna Seshadri for organizing this event smoothly. Also, a big thanks to Varun, Ezhilarasi, Ashish, Kumaresan, Rakesh, Kishore, etc. and the other folks in Chennai who helped out with the event. It was a big success!

Indian Sudoku Championship 2016


The Indian Sudoku Championship (ISC) 2016 was held on 16th July, 2016 in Chennai.

The championship was an offline event (after three years of having it online), where the top-50 qualifiers of Sudoku Mahabharat were invited. This is the first time the reigning champion did not defend the title. Rishi Puri, who won ISC in 2014 and 2015, has called it quits for his sudoku career.

View Top Qualifiers

Last 3yrs were very exciting competing with Prasanna and Rishi. Its unfortunate that neither of them participated this year since Prasanna, being the best Indian at the World Championships last year, gets a wild card for the national team this year and hence, decided to organize the national event.

I've traditionally done well in offline ISCs... in fact, only 3 of the ISCs were offline (2010, 2011, 2012), and those were the 3 times I won! This year I made it 4/4.


Puzzles
The championship consisted of four rounds. When you have Deb and Prasanna authoring sudokus, you know you're in for a treat. Really nice set of puzzles.

The fourth round of 6x6 sudokus had a fun twist to it, and was well thought of. The timings were perfectly set and the event ran very smoothly.

Those of you who'd like to solve these puzzles can purchase it here (Unfortunately due to certain reasons, they are not publicly available this year).


Results
I'm glad I topped all four rounds and won my fourth ISC title, this one after 4 years!

Rakesh Rai pipped Kishore Kumar for 2nd place as Kishore had a really bad 4th round. Most of the other results were as expected. Akash Doulani and Pranav Kamesh were the top inexperienced players comfortably, and I'm glad they'll finally be making it for their first World Championships.

View Complete Results

The winning trophies were really cool! Custom-made sudokus by Prasanna printed on the trophy with the shapes '1', '2' and '3'. Thanks to Sumit Bothra for getting these made.






Organization
Thanks to Deb Mohanty and Prasanna Seshadri for organizing this event smoothly. Also, a big thanks to Varun, Ezhilarasi, Ashish, Kumaresan, Rakesh, Kishore, etc. and the other folks in Chennai who helped out with the event. It was a big success!

Thursday, May 5, 2016

The Seer's Accuracy


AnalyticsVidhya organized a weekend hackathon The Seer's Accuracy on 29th April - 1st May, 2016.

In the midst of a new job, new city, I wasn't sure if I'll get enough time to participate in this hackathon. But fortunately, it was a relatively light weekend.

Problem
The challenge was to predict which customers would be return customers to a chain of stores. Looking at it another way, it was predicting which customers would churn.

Data
The train data consisted of customers (a.k.a. clients) and their transaction history in the years 2003 - 2006. The test evaluation was on which clients would return in 2007.

There was no test data per se, and it turned out to be the most crucial part of this challenge.

Overall, very clean data and very interesting problem. Kudos to AV!

Model
Right from the beginning I felt setting up a validation framework is going to be the key. And with a few LB submissions, I realized it was going to be extremely important to have a good validation set too.

I started off just like most other participants by using 2003-04-05 as build set and 2006 as the validation set and running a CV on it.
What finally catapulted me up the LB was when I added 2003-04 as build and 2005 as the validation set as well in my CV framework.

I think this resulted in a much stabler validation set and my CV and LB improvement was much more in sync.

Since the variables were limited, I treated and tested each of them individually and finally had a model with 335 features.

My final model was a blend of 3 XGBs on varying subsets of data and features.
It was a very minor improvement over my single best model.

GitHub
View My GitHub Repository

Results
I stood 1st on the public LB scoring 0.8856 and 1st on the private LB too, scoring 0.8800 using the AUC metric with the username 'vopani'.

Congrats to orenov/DataGeek for 2nd place and Bishwarup for 3rd place.

View Final Results

Views
Feels good. Really good.
Not just for winning, but for building a solid architecture which enabled a strong and stable model resulting in a considerable lead over the rest.

And this is also my first win on AV! :-)

Thanks to the AV organizers for this hackathon, was top quality and totally worth spending a weekend over.

View AV article on winners.

Sunday, March 6, 2016

Telstra Network Disruptions


The Telstra Network Disruptions competition was held on Kaggle in Nov, 2015 - Feb, 2016.

Objective
The objective was to predict the severity of a service disruption (whether it is a momentary glitch or a serious interruption of connectivity) on the Telstra network.

Data
The data consisted of disruptions along with features related to logs, events, resources, severity types across various locations.
The target variable was the severity of the disruption, into 3 classes.

Model
There was a golden insight in the data, and exploiting that became a very interesting challenge.

I ensembled several XGBoosts, on different subsets of the data and features, and some combinations of parameters.

The features used were the one-hot encoded raw features along with some interesting features built using the golden insight.

GitHub
View My GitHub Repository

Results
I stood 10th on the public LB and 9th on the private LB, scoring 0.40735 / 0.40267 using the logloss metric. My username is 'Vopani' and I competed as 'Anonymous Ghost' during the competition.

Views
This contest was all about finding and hacking that golden feature. The feature was nothing complex, it was simply the ordering of the observations that mattered. It was hard to spot because the ordering held true in the feature files (log, event, resource, etc.) and not on the original train/test data.

After identifying the relevance of ordering, it was very interesting to work on feature engineering and build features to improve the model without overfitting.

I'm glad I finished in the Top-10 and it also becomes my first Top-10 finish on Kaggle as an individual. I've moved to 70th in overall Kaggle rankings. I should easily be able to get into Top-50 by the end of the year. Maybe even Top-25.

Check out My Best Kaggle Performances

Saturday, February 13, 2016

AirBnB New User Bookings


The AirBnB New User Bookings competition was held on Kaggle in Nov-15 to Feb-16.

Objective
The objective was to predict in which country a new user on AirBnB would make their first booking.

There were 11 potential countries along with a 12th class - NDF (No Destination Found), indicating the user did not make any booking.

Data
The data consisted of user characteristics like language, age, browser, date-of-account-creation, OS, etc. for the train and test users.

There was data on the actions taken by users on the website along with the details of the action and duration.

Model
It was evident that the best way to quickly get a good score was to focus on classifying the NDF vs non-NDF users. So, I built a Logistic Regression on the one-hot encoded action features from the sessions data as a binary classifier for NDF vs non-NDF, only considering users present in the sessions data. This was the base classifier.

I then built a meta classifier using, well, everyone's favourite nowadays, XGBoost. It used the raw user features, along with the one-hot encoded features from sessions data, and finally, the LR predictions.

I did not complicate the model or ensemble too much due to lack of time, and also since the CV and LB were not perfectly correlating. Hence, I chose fairly simple models with some feature engineering.

GitHub
View GitHub Repository for the complete code, results and output.

Results
This model scored 0.88081 on the public LB which was ranked 89 and scored 0.88625 on the private LB which was ranked 23.
The metric used was NDCG.

View Public LB
View Final Results

Views
It was a very interesting dataset, and a good practise in building features from the sessions data, and without that, it wasn't possible to get a good score. It was disappointing that I had so many ideas which involved a lot more time to try out and code, but wasn't able to.

So, I think it was a simple stable model with lesser overfitting compared to many other competitors who dropped on the private LB.
In the end, I'm happy with the result, and this improves my overall Kaggle rank to 96th. So, finally I get into the Top-100 and on the first page of the rankings :-)

Hoping to improve on this further this year, and hopefully get into the Top-50 or Top-25 some day.

Check out My Best Kaggle Performances

Thursday, February 11, 2016

Puzzle Grand Prix 2016


The WPF Puzzle Grand Prix 2016 is here! After successful editions in 2014 and 2015, the 2016 edition consists of eight rounds held across the first seven months.
The top-10 finalists will be invited for the GP playoffs during WPC 2016 (Slovakia).

The format has changed a bit this year, with the contest having two sections. A Competitive section which is the main section on which the toppers will be decided, and a Casual section, comprising of more 'culture-neutral', non-grid based puzzles, geared towards leisure solvers. Its very unlikely competitors will be able to participate in both sections, so, I guess most players need to choose one.

I think I'm better at the 'Casual' sort of puzzles, and hence, will be competing only in that section.

Scoring System
There has been some discussions regarding the new scoring system for the GPs. Historically, normalization of scores for a championship consisting of various rounds has worked well due to the unreliability of having similar rounds in terms of scores and difficulty.
The GPs did use normalization in the previous editions, and it was universally accepted.

I'm not convinced there was a need to go ahead without normalization this year, so it remains to see whether or not it will work. You can find some pros/cons being discussed about the new scoring system on the GP Forum.
Being part of the organizing team of Sudoku Mahabharat / Puzzle Ramayan (which are very similar to the structure of GPs), we had to change the scoring system during this year's rounds, due to the inconsistencies without normalization. So, I'm not particularly in favour of dealing with raw scores.

View Championship Page
View Current Rankings



Round 6: Serbia (10th - 13th Jun, 2016)
As the competition is getting heated up, I was fairly comfortable with the puzzles in this round. Enjoyed the set and managed to finish all but one puzzle (Weights).

Before this round, I was leading with ~ 25points over Adam Bissett and ~ 50 points over Yuhei Kusui. Ironically, all three of us scored the exact same points: 344, in this round, keeping the standings the same.

It now boils down to the last two rounds to decide the winner.


Round 5: USA (13th - 16th May, 2016)
I was afraid that I barely managed to score 200 points in this hugely big 500+ pointer round. But it seems like it was too hard and everyone struggled.

The concept of Escape The Grand Prix is really nice, but not suitable for a time-contrained online puzzle round. This is easily going to be the discarded round for most of the top players.

Kudos to Randy Rogers for scoring 254 points, and Sinchai Rungsangrattanakul for scoring 242, way above the rest of the lot, while I scored just 205. Which is my poorest round so far.


Round 4: Hungary (15th - 18th Apr, 2016)
An easy set here. Managed to finish all and score full points, thus keeping my lead intact.
Nice puzzles too.

The four rounds so far have been worth 293, 402, 420 and 259 points. What in the world is the logic of not having normalization. That is the biggest failure of this year's GPs.


Round 3: Germany (18th - 21st Mar, 2016)
What an amazing set of puzzles. Just wow! This is the best set of puzzles I've solved in a long time... each and every one of them is a beauty.

The Instructionless Machine puzzles... exceptionally interesting and well-made. I think it was the sheer fun of the round that made me perform so well. I topped with 420 points, scoring over 100 points more than everyone else, except Jarett Prouse who scored 382.
Which means, I'm back in the top-3.


Round 2: Slovakia (19th - 22nd Feb, 2016)
Bad round. Lost time on the scrabble puzzles and wasn't able to complete them at the end.

Scored 283 points which is bad compared to the highest being 402. Hopefully, this could be one of my discards.

With no normalization, we have two rounds, one worth 293 points and the other worth 402 points. I just don't get it.


Round 1: India (22nd - 25th Jan, 2016)
It was a spur of the moment decision to participate in this round, since, being authored by Indians, I just assumed I couldn't compete. But Prasanna informed me that I was the test-solver only for the Competitive section and I could, in fact, participate in the Casual. And I did.

I finished 4th with 262 points. The highest score was 275 points by Adam Bissett of UK. Not a bad start.

The puzzles were excellent. The Buttons and the Number Series really got me scratching my head for a long time, and ultimately ended up missing out on two Number Series and a minor error in Shape Count.

Its quite funny that considering normalization is not being used this year, you'd expect the rounds to at least be similar in terms of points. The Competitive section was worth 697 points while the Casual section was worth 293 points. So, I have no clue where this is going. Lets hope its not too bad.

Wednesday, January 27, 2016

Sudoku Grand Prix 2016


The WPF Sudoku Grand Prix 2016 is here! After successful editions in 2014 and 2015, the 2016 edition consists of eight rounds held across the first seven months.
The top-10 finalists will be invited for the GP playoffs during WSC 2016 (Slovakia).

Due to certain unfortunate incidents, I wasn't able to compete completely in the first two editions of the GP. Well, I'm really hoping to make it this year.
Rishi Puri, the current and two-time national champion, featured in the playoffs in both years while Prasanna Seshadri was in the playoffs last year. With Rishi 'retiring' from active participation, it'll be interesting to see how the GP unfolds for the Indians!

Scoring System
There has been some discussions regarding the new scoring system for the GPs. Historically, normalization of scores for a championship consisting of various rounds has worked well due to the unreliability of having similar rounds in terms of scores and difficulty.
The GPs did use normalization in the previous editions, and it was universally accepted.

I'm not convinced there was a need to go ahead without normalization this year, so it remains to see whether or not it will work. You can find some pros/cons being discussed about the new scoring system on the GP Forum.
Being part of the organizing team of Sudoku Mahabharat / Puzzle Ramayan (which are very similar to the structure of GPs), we had to change the scoring system during this year's rounds, due to the inconsistencies without normalization. So, I'm not particularly in favour of dealing with raw scores.

View Championship Page


Round 3: Czech Republic (4th - 7th Mar, 2016)
Coming soon...


Round 2: Serbia (5th - 8th Feb, 2016)
Ohhh no. I bombed this round. Just had a bad day. And another submission mistake made it worse. So, that makes it two bad rounds out of two :-(

Puzzles were nice, nothing extraordinary. Many of them turned out to be Converse-like variants, which incidently comes right before the Converse round of SM, so good practise there.


Round 1: Netherlands (8th - 11th Jan, 2016)
Not the start I was hoping for. Had a rough solve, got stuck up time and again. To make things worse, I had one submission incorrect.

Tiit finished the set in just less than an hour, which is phenomenal. Prasanna finished the set in 79mins putting him in 5th. I finished in 84mins, but the mistake dropped me to 14th place.
The puzzle quality was excellent, as expected from the Dutch authors.

I hope this is the round that gets discarded!

Monday, January 4, 2016

Classic Tapa Contest 2016


Classic Tapa Contest (CTC) 2016 was held in Jan-Feb, 2016 on LMI.

Championship Page

View Forum
View Results

1. sai (Japan)
2. EKBM (Japan)
3. Prasanna16391 (India)
4. deu (Japan)
5. Psyho (Poland)
6. Para (Netherlands)
7. nyoroppyi (Japan)
8. willwc (USA)
9. kiwijam (New Zealand)
10. anderson (USA)

View Complete Results

sai won the last CTC and this time completely dominated at the top. After a horrific start, Endo blew through the middle days and just managed to overtake Prasanna in the last week to finish 2nd. And phenomenal performance by Prasanna, who takes 3rd, after an extremely consistent run of two months.

I was in the race to get into the Top-20, but unfortunately missed a few days due to personal matters. Nevertheless, I'm quite happy with my performance and enjoyed the Tapas.

I'm sure a lot of people are going to miss CTC... it just gets on you after 50 days :-)