Search This Blog

Monday, February 26, 2018

Moving to Medium

I will be sharing my writings henceforth on Medium. Here's my handle: https://medium.com/@vopani

See you there!

Sunday, February 5, 2017

Blackjack Numberlink


Nibbl is a puzzle solving and publishing mobile platform where users can solve hand-crafted puzzles created by some of the top puzzlers of the world.

As a follow-up to my Blackjack puzzle series and my previous Nibbl collection of Hitori puzzles, I've published my second collection called 'Blackjack Numberlink'. Its contents are:

6x6 Numberlink - 2 puzzles
7x7 Numberlink - 2 puzzles
8x8 Numberlink - 6 puzzles
9x9 Numberlink - 4 puzzles
10x10 Numberlink - 7 puzzles

If you enjoy solving Numberlink puzzles, or want to try your hand on a few, you can purchase this from Nibbl and solve it on your phone at leisure!

Currently, the collection with these 21 puzzles is priced at 99 Nbls (little less than $1). Not only can you solve them, but you can also see the solution path set by the author to help you understand the flow of the solution. Maybe learn some tricks with it too.

You'll find it in the Collections section in Nibbl. These are not published as individual puzzles, so, its only possible to solve them by purchasing the collection.

Check out some other interesting collections by fellow Nibbl authors: Rob Vollmert's Star Battle and Prasanna Seshadri's Tapa collections are worth a shot.

For those of you who aren't on Nibbl yet, you can download the app from Apple Store or Google Play Store. Feel free to use my referral code for some bonus credits: ROHA3286 Hope you enjoy the set!

I plan to publish similar collections like these in the months to come. A Sudoku collection will be published soon!

Monday, November 28, 2016

Blackjack Hitori


Earlier this year, I posted about Nibbl, a puzzle solving and publishing mobile platform where users can solve hand-crafted puzzles created by some of the top puzzlers of the world.

As a series of packs (or 'Collections' on Nibbl), I've published my first set called 'Blackjack Hitori'. Its contents are:

4x4 Hitori - 1 puzzle
5x5 Hitori - 1 puzzle
6x6 Hitori - 1 puzzle
7x7 Hitori - 1 puzzle
8x8 Hitori - 5 puzzles
9x9 Hitori - 5 puzzles
10x10 Hitori - 7 puzzles

I tried to create these with varying difficulty levels and as many different types of solving paths as possible. Hopefully, with some pleasant themes in some of them :-)

If you enjoy solving Hitori puzzles, or want to try your hand on a few, you can purchase this from Nibbl and solve it on your phone at leisure!

Currently, the collection with these 21 puzzles is priced at 99 Nbls (little less than $1). Not only can you solve them, but you can also see the solution path set by the author to help you understand the flow of the solution. Maybe some of you can learn some tricks with it too.

You'll find it in the Collections section in Nibbl. These are not published as individual puzzles, so, you can only solve them from the whole collection.

You can also check out some interesting collections by fellow Nibbl authors. Rob Vollmert's Star Battle and Prasanna Seshadri's Tapa collections are worth a shot.

For those of you who aren't on Nibbl yet, you can download the app from Apple Store or Google Play Store.

Feel free to use my referral code for some bonus credits: ROHA3286

Hope you enjoy the set!

I plan to publish similar packs like these in the months to come. A Sudoku collection will be published next month, so watch out for it.

Special thanks to Dan Adams, Rob Vollmert and Prasanna Seshadri for testing the puzzles.

Thursday, November 3, 2016

World Puzzle Championship 2016


The 25th World Puzzle Championship (WPC) was held on 20th - 22nd Oct, 2016 in Senec, Slovakia.

View Championship Page


Indian Team









The Indian team was selected from the Indian Puzzle Championship 2016 (IPC). We had two complete teams this year.

The A-Team is same as last year's.


Preparation
Similar to my WSC experience, I didn't have too much preparation or expectation this year. As a team, we were hoping to get into a single digit rank, which we missed last year with our 10th place.


Puzzles
Similar to my WSC experience, puzzles were fantastic all around. Some beautiful themes and well-crafted puzzles.


Performance
The Indian team performed well this year. Prasanna improved on his (and India's best individual rank) by finishing 18th (one better than last year's 19th), which is also his fourth consecutive finish in the Top-25. It is quite creditable considering no other Indian has ever crossed the 25th mark.

I finished 28th, which is my best WPC rank ever! Yoohoo! 8yrs later, when I'm old and on the verge of quitting, comes my best performance. Haha, well, it certainly felt good but I'm not sure if I'll ever be able to cross the 25th mark.

I made a single digit error in two high-pointers, one of which was the Full Scrabble. Had I got that Scrabble right, it would be the first WPC round I'd have convincingly finished (ignoring the single puzzle rounds). Well... maybe some other time.

Amit finished 39th and Swaroop, a disappointing 72nd.

We finished 10th in the team standing, equalling our best team rank last year. Missed single digit yet again.

The unofficial players: Rakesh did well in 69th, Ashish in 108th, Rajesh in 119th, Kishore in 133rd and Jaipal in 151st.

Prasanna has been in tremendous form in the GPs, but had an average WPC this year.
He's doing well in elections too! (Yes, he got elected as a board member in the World Puzzle Federation, but more on that later)


Winners
Last year, we saw Ken Endo (Japan) pip Ulrich Voigt (Germany) in the playoffs. This year Ulrich pipped Endo to take his 11th WPC title. Absolutely incredible, and he still continues to be on top after all these years.

Endo had a huge lead in the preliminary rounds, over Ulrich and Palmer. Scoring 1000+ points over Ulrich is truly a great achievement. It felt like the start of Endo's domination after his string of magnificent performances in online contests throughout the year. But it was not to be...

I missed the playoffs due to a Machine Learning hackathon I was competing in, but Endo unfortunately fell to third, with Palmer finishing 2nd.

View Complete Results


Puzzle GP
I've loved the GPs and its high-quality puzzles, and I've shared most of my thoughts on my WSC blogpost.

This was the first time I competed in Puzzle GP, only because I really enjoy the Casual-type puzzles. After the first two rounds, it was evident there wasn't much competition here (well, the Competitive section is what really matters), but I thought I'll go ahead and compete for the Casual winner.

And I won! There were no playoffs for the Casual section, but the Competitive section playoffs was quite entertaining.

Endo raced through the puzzles for an easy win, but it was a fight to the finish for 2nd and 3rd, with 7 of the 10 players on the last puzzle. Ulrich finished 2nd, just 3 seconds before Will Blatt (USA).

View Casual GP Results
View Competitive GP Results
View General GP Results


Organization
Kudos to Matus Demiger, Matej Uher, Peter Hudak and the entire Slovak team for hosting this wonderful event. It was excellently organised and a lot of fun! Thanks a lot!

There were far too many Kropki puzzles (ask Przemyslaw about it), but well, every team have their preferences.


Next Year
Slovakia have set the bar high, and the Indian team will be hosting all you wonderful people next year during the same week (15th - 22nd Oct, 2017) in Bengaluru, India. I look forward to it as much as many of you, and hope it is a successful and enjoyable WPC.

Feel free to write to me: rohanrao88@gmail.com in case of any help/queries/questions regarding next year's event. We'll have our webpage out soon, and the LMI forum is always available and active for other discussions!

And now, back to boring non-puzzle life :-|

P.S. Read about my WSC experience.

Wednesday, October 26, 2016

World Sudoku Championship 2016


The 11th World Sudoku Championship (WSC) was held on 17th - 18th Oct, 2016 in Senec, Slovakia.

View Championship Page


Indian Team











The Indian team was selected from the Indian Sudoku Championship 2016 (ISC) and the Times Sudoku Championship 2016 (TSC). We had three complete teams this year.

Pranav was our Under-18 participant, and after a good show at last year's World Junior Championships, he made it to the Top-2 inexperienced solvers of TSC this year.


Preparation
I've been off the grid, primarily due to my work and career aspirations of pursuing Machine Learning, so I didn't have high expectations this year. This might also be my last World Championships as a participant, as I'm one of the organizers for next year's championships in India, after which I'm uncertain of continuing.


Puzzles
The Slovak team have been doing well in recent years, in contests, in authoring, in communicating, and I think a lot of people were expecting these championships to be good fun. And it definitely was. The IB was crisp, rounds were well thought of, themes were really nice.

Retrospectively, the timing of all rounds were spot on, and it was one of my best WSC experiences in the last 8 years.

The puzzles were all well-made, a good mix of easy, medium and hard puzzles, a good variety of standard vs new variants, some really exciting team rounds, and overall it felt like a complete well-balanced WSC. Thumbs up for all the authors and organizers for the excellent event.

I must admit that there were some parts where I felt deja-vu, strikingly similar to some of the ideas we've been discussing for next year's WSC. Well, there's no limit to innovation!


Performance
The Indian team performed average-ish this year. I was the top ranked Indian on 18th, which is our worst 'top rank' in the last 9 years! But as a team, all four of us were ranked in Top-40, which is a first.
Prasanna finished 21st, Rakesh 32nd and Kishore 38th. Rakesh's and Kishore's personal best.

We finished 6th in the team standing, equalling our best team rank in 2014.

The unofficial players: Amit stood 53rd, Gaurav 73rd, Rajesh 76th, Swaroop 85th, Jaipal 93rd, Pranav 94th, Akash 104th and Ashish 141st.

Prasanna has been in tremendous form in the GPs, but had quite a bad WSC this year.
He's doing well in elections too! (Yes, he got elected as a board member in the World Puzzle Federation, but more on that later)


Winners
Jakub Ondrousek (Czech Republic) performed extremely well in the rounds and was on top with a sizeable lead over Tiit Vunk (Estonia) and Kota Morinishi (Japan). The three of them have been consistently topping the WSCs last few years, and this one was no different.

But it was the Estonian who flew through the playoff puzzles and pipped Jakub to take his first title, after finishing runner-up in the previous two years. Hearty congratulations to Tiit! A well-deserved win.
Jakub was second and Kota stood third.

Among the teams, Czech Republic blazed through the playoffs to win it, while the Chinese came in second, with an amazing performance, where three of them were in Top-10 too (let aside the fact that they are all U-18!) and Japan came in third.

View Complete Results


Sudoku GP
I've always enjoyed the GP rounds through the year, but never perform well in them, quite unlike the WSCs. Somehow, it makes me feel I'm more of an offline solver, rather than an online one (Even of the 10 ISCs so far, only 4 were held offline, and those were the 4 I won!).

There was a lot of debate on the new scoring system, and I didn't share my thoughts earlier since I was just uninterested and busy. But, I ended up 14th in the GP, and due to high dropouts in the Top-10, I was invited for the playoffs.

And I decided not to compete. There are a couple of reasons for this. One is, I reviewed my performance and felt I didn't deserve to be in the playoffs. Using some simple normalization, I would lose at least 2 more ranks, which is a fairer position for me. And secondly, the schedule this year had more rounds, and having the GP in between them is just exhausting. Besides the fact that I wanted a decent rank in the WSC if this is my last one, so preferred resting.

I don't want to comment too much on the scoring, but there's no harm in trying new ideas, I'm always in favour of experimenting. Tom Collyer, the Sudoku GP Director, received some slack on this, and I'm not sure its completely fair to him.

It is just my personal opinion and preference that some form of normalization does make a better scoring system. We (LMI organizers) have seen it in the past... we even changed it in the middle of SM and PR, and decided to add normalization after 2 rounds, when we realized it is extremely difficult to curate 8 rounds of 'equal' solving experience.

It remains to see what happens next year. The Sudoku GP Director spot is currently open (Tom stepped down).

Overall, the GPs were still a lot of fun with high-quality puzzles. Thanks to Tom, Hana, Wei-Hwa, all the authors, testers and everyone else who put in effort conducting these year-on-year.

The top-3 in the playoffs were Tiit Vunk (so he won the double), Kota Morinishi and Hideaki Jo. Congrats!


Organization
Kudos to Zuzana Hromcova and the entire Slovak team for hosting this wonderful event. It was smooth, it was punctual, it was lively, it was fun! Thanks a lot!

This year, there were several prizes (and some spot prizes too!), which is generally nice to have.

There was a prize for guessing the 11th place player (being the 11th WSC), and we analyzed the results after two rounds, to choose the most probable twelve people who'd be 11th, distributing the votes among us twelve Indians.
And we bombed it like anything. We ended up guessing the players ranked 9th, 10th, 12th, 13th, 14th, 15th, 16th, 18th, 19th, etc. Damn we missed Takuya! :-)


Next Year
Slovakia have set the bar high, and the Indian team will be hosting all you wonderful people next year during the same week (15th - 22nd Oct, 2017) in Bengaluru, India. I look forward to it as much as many of you, and hope it is a successful and enjoyable WSC.

I've got some plans through the year, so, its going to be an interesting next few months and I hope most of them work out as expected. Fingers crossed!

Feel free to write to me: rohanrao88@gmail.com in case of any help/queries/questions regarding next year's event. We'll have our webpage out soon, and the LMI forum is always available and active for other discussions!

And now, back to boring non-puzzle life :-|

P.S. Will be writing an update on WPC by the weekend.

Wednesday, August 3, 2016

The Smart Recruit


AnalyticsVidhya organised a weekend hackathon called The Smart Recruit, which was held on 23rd-24th July, 2016.

I won the previous hackathon, The Seer's Accuracy, and was hoping to do well in this one too.

Problem
The problem was to identify which agents would be successful in making sales of financial products. So, it was a binary classification problem.

Data
Like the previous hackathon, the data seemed quite good and promising.

Train and test data consisted of agent applications with data about the application, the manager and a few related features about them.

Model
I'm sure most participants just went ahead and dumped the data into XGB-type models with a lot of scores hovering in the 0.63 - 0.66 AUC range.

I tried to get a robust/stable validation framework, like I mentioned in AV's article on Winning Tips. Didn't seem to work/help. The CV/LB scores were all over the place in my first few submissions.

Thats when I decided to take a step back and inspect the data in detail. It was evident to me that there would be a huge LB shake-up due to the variance between the CV-LB scores. Hence, didn't make much sense to spend too much time on the data trying to optimize models. Instead, I tried to look for some pattern/feature which could boost me score over the expected error margins.

And that's exactly what happened. A simple plot of the target variable showed a pattern, which seemed too good to be true. I tried a feature using this and my CV jumped to 0.8... and that was the feature that ultimately proved to be the winning one.

Here's the plot that changed everything:

This is the plot of the target variable for the first four days. A clear pattern exists where you see most of the 1's at the beginning of a day and most of the 0's at the end of the day. You can plot the target variable of any single day and observe a similar trend.

Leakage? Possible. Hidden trend? Possible. At first I was convinced it was leakage and a data preparation issue, but later, felt there was a possibility that applications received towards the end of the day are more likely to be rejected than ones received early.

Either ways, I polished this feature using Order_Percentile in my code, which was the most important feature.

My final model was a single XGBoost with 14 features, with the other 13 being cleaned up features from the raw variables. I achieved a CV of 0.887 which was in the same range as the LB. I'd have liked to try out some more parameter tuning and ensembling, but with the limited duration of a hackathon, there wasn't any time left.

GitHub
View My Complete Solution

Results
I stood 1st on the public LB with 0.885, with good friend and rival competitor SRK in 2nd, who teamed up with Kaggler Mark Landry, with 0.876 and another team of Kanishk Agarwal and Yaasna Dua in 3rd with 0.839. No other team figured out the winning feature and their scores were below 0.71.

The rankings held same on the private LB, but it was much closer, with SRK-Mark scoring 0.7647 and I scoring 0.7658.

My username is 'vopani'.

View Complete Results

Views
My 2nd AV win on the trot and while not the best way to win it, I'm happy I could find a useful winning feature in the data.

Congrats to the ever consistent SRK, who also happens to be someone I'm chasing on Kaggle :-)

Fun weekend, bonus to win it, and looking forward to the next hackathon, where I'll be on a hat-trick!

An interesting co-incidence: I got the exact same score on the public LB (0.8856) in the previous hackathon too, The Seer's Accuracy !!!

External Links
View AV article on the winners
View 2nd place solution by SRK
View 3rd place solution by Kanishk Agarwal

Friday, July 29, 2016

Nibbl


Nibbl is a platform for authoring and solving various grid-based logical puzzles. Being a part of the puzzling world, having authored, organized and participated in various international puzzle championships across the globe, I've wondered what is the best way to reach out, publicise and improvise on these popular genres across different channels and audiences. With the advancements of mobile technology, it is one of the fastest and most convenient forms of information distribution.

So, why Nibbl?

Yes, there are a few puzzle solving (especially sudoku) apps out there, but most of them are computer generated puzzles built by programmers. Don't we all love hand-crafted theme-based puzzles with wonderful solving paths to 'tickle our minds'? All puzzles on Nibbl are custom created by authors.

One of the best features in Nibbl is the ability to view the author's intended solving path. A great way to learn from the creators on what logic was applied when and how.

Currently, it has a few popular puzzle genres but more will be added soon.
Nibbl Authors can publish puzzles as per their choice and liking and decide their own pricing. Creating puzzles is a matter of few taps and can be done on the app itself. Interested authors can contact nibble.appfactory@gmail.com.

Solvers can download the app from:
Google Play Store: https://play.google.com/store/apps/details?id=rao.rajnikant.ips
Apple Store: https://itunes.apple.com/in/app/nibbl-solvr/id1132803199

Feel free to use my referral code for some bonus credits: ROHA3286

I plan to publish some puzzle packs on Nibbl later this year.

P.S. My father, Rajnikant Rao, is the main driver of the project :-)

Friday, June 24, 2016

Indian Puzzle Championship 2016


The Indian Puzzle Championship (IPC) 2016 was held on 17th July, 2016 in Chennai.

The championship was an offline event (after four years of having it online), where the top-50 qualifiers of Puzzle Ramayan were invited. Prasanna is undoubtedly the best puzzle solver in India today, and has been last 3yrs.

View Top Qualifiers

Last 2yrs were exciting competing with Amit. He won the 2014 edition and I came in second, while I won the 2015 edition and he came in second. I was expecting him to be the biggest challenger for the title, since Prasanna, being the best Indian at the World Puzzle Championship last year, gets a wild card for the national team this year and hence, decided to organize the national event.


Puzzles
The championship consisted of four rounds. When you have Deb and Prasanna authoring puzzles, you know you're in for a treat. Really nice set of puzzles.

The 3rd round Sprint was my favourite, and also, the only round I topped.
The 4th round was Casual-type puzzles, many of them visual, and supposedly one of my strengths. But, I did horribly.

Those of you who'd like to solve these puzzles can purchase it here (Unfortunately due to certain reasons, they are not publicly available this year).


Results
I expected it to be a close fight between Amit and me. But, after the first two rounds, Amit had a big enough lead over me and it was pretty hard to catch up from then. I did cover up a few points in the third round, but had a terrible fourth round.

Amit won his second IPC and I stood 2nd (for the third time, after 2009 and 2014).

The ever-consistent Rakesh stood 3rd with a strong performance. Ashish's performance was disappointing, but just made it into the B-Team with his 6th place finish.

View Complete Results

The winning trophies were really cool! Custom-made puzzles by Prasanna (and one by me, the second one, which I ended up getting :-| ) printed on the trophy with the themes '1', '2' and '3'. Thanks to Sumit Bothra for getting these made.




Organization
Thanks to Deb Mohanty and Prasanna Seshadri for organizing this event smoothly. Also, a big thanks to Varun, Ezhilarasi, Ashish, Kumaresan, Rakesh, Kishore, etc. and the other folks in Chennai who helped out with the event. It was a big success!

Indian Sudoku Championship 2016


The Indian Sudoku Championship (ISC) 2016 was held on 16th July, 2016 in Chennai.

The championship was an offline event (after three years of having it online), where the top-50 qualifiers of Sudoku Mahabharat were invited. This is the first time the reigning champion did not defend the title. Rishi Puri, who won ISC in 2014 and 2015, has called it quits for his sudoku career.

View Top Qualifiers

Last 3yrs were very exciting competing with Prasanna and Rishi. Its unfortunate that neither of them participated this year since Prasanna, being the best Indian at the World Championships last year, gets a wild card for the national team this year and hence, decided to organize the national event.

I've traditionally done well in offline ISCs... in fact, only 3 of the ISCs were offline (2010, 2011, 2012), and those were the 3 times I won! This year I made it 4/4.


Puzzles
The championship consisted of four rounds. When you have Deb and Prasanna authoring sudokus, you know you're in for a treat. Really nice set of puzzles.

The fourth round of 6x6 sudokus had a fun twist to it, and was well thought of. The timings were perfectly set and the event ran very smoothly.

Those of you who'd like to solve these puzzles can purchase it here (Unfortunately due to certain reasons, they are not publicly available this year).


Results
I'm glad I topped all four rounds and won my fourth ISC title, this one after 4 years!

Rakesh Rai pipped Kishore Kumar for 2nd place as Kishore had a really bad 4th round. Most of the other results were as expected. Akash Doulani and Pranav Kamesh were the top inexperienced players comfortably, and I'm glad they'll finally be making it for their first World Championships.

View Complete Results

The winning trophies were really cool! Custom-made sudokus by Prasanna printed on the trophy with the shapes '1', '2' and '3'. Thanks to Sumit Bothra for getting these made.






Organization
Thanks to Deb Mohanty and Prasanna Seshadri for organizing this event smoothly. Also, a big thanks to Varun, Ezhilarasi, Ashish, Kumaresan, Rakesh, Kishore, etc. and the other folks in Chennai who helped out with the event. It was a big success!

Thursday, May 5, 2016

The Seer's Accuracy


AnalyticsVidhya organized a weekend hackathon The Seer's Accuracy on 29th April - 1st May, 2016.

In the midst of a new job, new city, I wasn't sure if I'll get enough time to participate in this hackathon. But fortunately, it was a relatively light weekend.

Problem
The challenge was to predict which customers would be return customers to a chain of stores. Looking at it another way, it was predicting which customers would churn.

Data
The train data consisted of customers (a.k.a. clients) and their transaction history in the years 2003 - 2006. The test evaluation was on which clients would return in 2007.

There was no test data per se, and it turned out to be the most crucial part of this challenge.

Overall, very clean data and very interesting problem. Kudos to AV!

Model
Right from the beginning I felt setting up a validation framework is going to be the key. And with a few LB submissions, I realized it was going to be extremely important to have a good validation set too.

I started off just like most other participants by using 2003-04-05 as build set and 2006 as the validation set and running a CV on it.
What finally catapulted me up the LB was when I added 2003-04 as build and 2005 as the validation set as well in my CV framework.

I think this resulted in a much stabler validation set and my CV and LB improvement was much more in sync.

Since the variables were limited, I treated and tested each of them individually and finally had a model with 335 features.

My final model was a blend of 3 XGBs on varying subsets of data and features.
It was a very minor improvement over my single best model.

GitHub
View My GitHub Repository

Results
I stood 1st on the public LB scoring 0.8856 and 1st on the private LB too, scoring 0.8800 using the AUC metric with the username 'vopani'.

Congrats to orenov/DataGeek for 2nd place and Bishwarup for 3rd place.

View Final Results

Views
Feels good. Really good.
Not just for winning, but for building a solid architecture which enabled a strong and stable model resulting in a considerable lead over the rest.

And this is also my first win on AV! :-)

Thanks to the AV organizers for this hackathon, was top quality and totally worth spending a weekend over.

View AV article on winners.

Sunday, March 6, 2016

Telstra Network Disruptions


The Telstra Network Disruptions competition was held on Kaggle in Nov, 2015 - Feb, 2016.

Objective
The objective was to predict the severity of a service disruption (whether it is a momentary glitch or a serious interruption of connectivity) on the Telstra network.

Data
The data consisted of disruptions along with features related to logs, events, resources, severity types across various locations.
The target variable was the severity of the disruption, into 3 classes.

Model
There was a golden insight in the data, and exploiting that became a very interesting challenge.

I ensembled several XGBoosts, on different subsets of the data and features, and some combinations of parameters.

The features used were the one-hot encoded raw features along with some interesting features built using the golden insight.

GitHub
View My GitHub Repository

Results
I stood 10th on the public LB and 9th on the private LB, scoring 0.40735 / 0.40267 using the logloss metric. My username is 'Vopani' and I competed as 'Anonymous Ghost' during the competition.

Views
This contest was all about finding and hacking that golden feature. The feature was nothing complex, it was simply the ordering of the observations that mattered. It was hard to spot because the ordering held true in the feature files (log, event, resource, etc.) and not on the original train/test data.

After identifying the relevance of ordering, it was very interesting to work on feature engineering and build features to improve the model without overfitting.

I'm glad I finished in the Top-10 and it also becomes my first Top-10 finish on Kaggle as an individual. I've moved to 70th in overall Kaggle rankings. I should easily be able to get into Top-50 by the end of the year. Maybe even Top-25.

Check out My Best Kaggle Performances

Saturday, February 13, 2016

AirBnB New User Bookings


The AirBnB New User Bookings competition was held on Kaggle in Nov-15 to Feb-16.

Objective
The objective was to predict in which country a new user on AirBnB would make their first booking.

There were 11 potential countries along with a 12th class - NDF (No Destination Found), indicating the user did not make any booking.

Data
The data consisted of user characteristics like language, age, browser, date-of-account-creation, OS, etc. for the train and test users.

There was data on the actions taken by users on the website along with the details of the action and duration.

Model
It was evident that the best way to quickly get a good score was to focus on classifying the NDF vs non-NDF users. So, I built a Logistic Regression on the one-hot encoded action features from the sessions data as a binary classifier for NDF vs non-NDF, only considering users present in the sessions data. This was the base classifier.

I then built a meta classifier using, well, everyone's favourite nowadays, XGBoost. It used the raw user features, along with the one-hot encoded features from sessions data, and finally, the LR predictions.

I did not complicate the model or ensemble too much due to lack of time, and also since the CV and LB were not perfectly correlating. Hence, I chose fairly simple models with some feature engineering.

GitHub
View GitHub Repository for the complete code, results and output.

Results
This model scored 0.88081 on the public LB which was ranked 89 and scored 0.88625 on the private LB which was ranked 23.
The metric used was NDCG.

View Public LB
View Final Results

Views
It was a very interesting dataset, and a good practise in building features from the sessions data, and without that, it wasn't possible to get a good score. It was disappointing that I had so many ideas which involved a lot more time to try out and code, but wasn't able to.

So, I think it was a simple stable model with lesser overfitting compared to many other competitors who dropped on the private LB.
In the end, I'm happy with the result, and this improves my overall Kaggle rank to 96th. So, finally I get into the Top-100 and on the first page of the rankings :-)

Hoping to improve on this further this year, and hopefully get into the Top-50 or Top-25 some day.

Check out My Best Kaggle Performances

Thursday, February 11, 2016

Puzzle Grand Prix 2016


The WPF Puzzle Grand Prix 2016 is here! After successful editions in 2014 and 2015, the 2016 edition consists of eight rounds held across the first seven months.
The top-10 finalists will be invited for the GP playoffs during WPC 2016 (Slovakia).

The format has changed a bit this year, with the contest having two sections. A Competitive section which is the main section on which the toppers will be decided, and a Casual section, comprising of more 'culture-neutral', non-grid based puzzles, geared towards leisure solvers. Its very unlikely competitors will be able to participate in both sections, so, I guess most players need to choose one.

I think I'm better at the 'Casual' sort of puzzles, and hence, will be competing only in that section.

Scoring System
There has been some discussions regarding the new scoring system for the GPs. Historically, normalization of scores for a championship consisting of various rounds has worked well due to the unreliability of having similar rounds in terms of scores and difficulty.
The GPs did use normalization in the previous editions, and it was universally accepted.

I'm not convinced there was a need to go ahead without normalization this year, so it remains to see whether or not it will work. You can find some pros/cons being discussed about the new scoring system on the GP Forum.
Being part of the organizing team of Sudoku Mahabharat / Puzzle Ramayan (which are very similar to the structure of GPs), we had to change the scoring system during this year's rounds, due to the inconsistencies without normalization. So, I'm not particularly in favour of dealing with raw scores.

View Championship Page
View Current Rankings



Round 6: Serbia (10th - 13th Jun, 2016)
As the competition is getting heated up, I was fairly comfortable with the puzzles in this round. Enjoyed the set and managed to finish all but one puzzle (Weights).

Before this round, I was leading with ~ 25points over Adam Bissett and ~ 50 points over Yuhei Kusui. Ironically, all three of us scored the exact same points: 344, in this round, keeping the standings the same.

It now boils down to the last two rounds to decide the winner.


Round 5: USA (13th - 16th May, 2016)
I was afraid that I barely managed to score 200 points in this hugely big 500+ pointer round. But it seems like it was too hard and everyone struggled.

The concept of Escape The Grand Prix is really nice, but not suitable for a time-contrained online puzzle round. This is easily going to be the discarded round for most of the top players.

Kudos to Randy Rogers for scoring 254 points, and Sinchai Rungsangrattanakul for scoring 242, way above the rest of the lot, while I scored just 205. Which is my poorest round so far.


Round 4: Hungary (15th - 18th Apr, 2016)
An easy set here. Managed to finish all and score full points, thus keeping my lead intact.
Nice puzzles too.

The four rounds so far have been worth 293, 402, 420 and 259 points. What in the world is the logic of not having normalization. That is the biggest failure of this year's GPs.


Round 3: Germany (18th - 21st Mar, 2016)
What an amazing set of puzzles. Just wow! This is the best set of puzzles I've solved in a long time... each and every one of them is a beauty.

The Instructionless Machine puzzles... exceptionally interesting and well-made. I think it was the sheer fun of the round that made me perform so well. I topped with 420 points, scoring over 100 points more than everyone else, except Jarett Prouse who scored 382.
Which means, I'm back in the top-3.


Round 2: Slovakia (19th - 22nd Feb, 2016)
Bad round. Lost time on the scrabble puzzles and wasn't able to complete them at the end.

Scored 283 points which is bad compared to the highest being 402. Hopefully, this could be one of my discards.

With no normalization, we have two rounds, one worth 293 points and the other worth 402 points. I just don't get it.


Round 1: India (22nd - 25th Jan, 2016)
It was a spur of the moment decision to participate in this round, since, being authored by Indians, I just assumed I couldn't compete. But Prasanna informed me that I was the test-solver only for the Competitive section and I could, in fact, participate in the Casual. And I did.

I finished 4th with 262 points. The highest score was 275 points by Adam Bissett of UK. Not a bad start.

The puzzles were excellent. The Buttons and the Number Series really got me scratching my head for a long time, and ultimately ended up missing out on two Number Series and a minor error in Shape Count.

Its quite funny that considering normalization is not being used this year, you'd expect the rounds to at least be similar in terms of points. The Competitive section was worth 697 points while the Casual section was worth 293 points. So, I have no clue where this is going. Lets hope its not too bad.

Wednesday, January 27, 2016

Sudoku Grand Prix 2016


The WPF Sudoku Grand Prix 2016 is here! After successful editions in 2014 and 2015, the 2016 edition consists of eight rounds held across the first seven months.
The top-10 finalists will be invited for the GP playoffs during WSC 2016 (Slovakia).

Due to certain unfortunate incidents, I wasn't able to compete completely in the first two editions of the GP. Well, I'm really hoping to make it this year.
Rishi Puri, the current and two-time national champion, featured in the playoffs in both years while Prasanna Seshadri was in the playoffs last year. With Rishi 'retiring' from active participation, it'll be interesting to see how the GP unfolds for the Indians!

Scoring System
There has been some discussions regarding the new scoring system for the GPs. Historically, normalization of scores for a championship consisting of various rounds has worked well due to the unreliability of having similar rounds in terms of scores and difficulty.
The GPs did use normalization in the previous editions, and it was universally accepted.

I'm not convinced there was a need to go ahead without normalization this year, so it remains to see whether or not it will work. You can find some pros/cons being discussed about the new scoring system on the GP Forum.
Being part of the organizing team of Sudoku Mahabharat / Puzzle Ramayan (which are very similar to the structure of GPs), we had to change the scoring system during this year's rounds, due to the inconsistencies without normalization. So, I'm not particularly in favour of dealing with raw scores.

View Championship Page


Round 3: Czech Republic (4th - 7th Mar, 2016)
Coming soon...


Round 2: Serbia (5th - 8th Feb, 2016)
Ohhh no. I bombed this round. Just had a bad day. And another submission mistake made it worse. So, that makes it two bad rounds out of two :-(

Puzzles were nice, nothing extraordinary. Many of them turned out to be Converse-like variants, which incidently comes right before the Converse round of SM, so good practise there.


Round 1: Netherlands (8th - 11th Jan, 2016)
Not the start I was hoping for. Had a rough solve, got stuck up time and again. To make things worse, I had one submission incorrect.

Tiit finished the set in just less than an hour, which is phenomenal. Prasanna finished the set in 79mins putting him in 5th. I finished in 84mins, but the mistake dropped me to 14th place.
The puzzle quality was excellent, as expected from the Dutch authors.

I hope this is the round that gets discarded!

Monday, January 4, 2016

Classic Tapa Contest 2016


Classic Tapa Contest (CTC) 2016 was held in Jan-Feb, 2016 on LMI.

Championship Page

View Forum
View Results

1. sai (Japan)
2. EKBM (Japan)
3. Prasanna16391 (India)
4. deu (Japan)
5. Psyho (Poland)
6. Para (Netherlands)
7. nyoroppyi (Japan)
8. willwc (USA)
9. kiwijam (New Zealand)
10. anderson (USA)

View Complete Results

sai won the last CTC and this time completely dominated at the top. After a horrific start, Endo blew through the middle days and just managed to overtake Prasanna in the last week to finish 2nd. And phenomenal performance by Prasanna, who takes 3rd, after an extremely consistent run of two months.

I was in the race to get into the Top-20, but unfortunately missed a few days due to personal matters. Nevertheless, I'm quite happy with my performance and enjoyed the Tapas.

I'm sure a lot of people are going to miss CTC... it just gets on you after 50 days :-)

Wednesday, December 16, 2015

Rossmann Store Sales


The Rossmann Store Sales competition was held on Kaggle in Nov-Dec, 2015.

Objective
Rossmann operates over 3000 drug stores in 7 European countries. Currently, Rossmann store managers are tasked with predicting their daily sales for up to six weeks in advance. Store sales are influenced by many factors, including promotions, competition, school and state holidays, seasonality, and locality. With thousands of individual managers predicting sales based on their unique circumstances, the accuracy of results can be quite varied.

The objective was to predict the sales of various Rossmann stores in Germany.

Data
Train data consisted of sales from over 1000 Rossmann stores along with information related to promotions, competitions, holidays, etc. upto July, 2015.

Test data consisted of dates in August and September, 2015 for which we had to predict the sales.

Approach
Being a classic time series sales forecasting problem, I explored two approaches. One being the standard building of tree-based and linear models. The other being trying out time series models like ARIMA.

It became quickly evident from cross-validation and validation results that ARIMA wasn't working. XGBoost was giving much better results.

There was a lot of external data shared and available, but none of those made a big improvement in the model. My final model didn't use any external data either.

Building models at a store-level was not giving as good results as building a model using all the data together, but it helped while blending models.

Model
I built multiple XGBoost models on different subsets of the entire data and averaged them. I merged these with store-level models of XGBoost, Random Forest and GBM. The blending of models gave a huge improvement and ultimately lead to the stability of the predictions.

I finally tweaked the predictions by using a multiplicative factor of 0.98 to get the best fit to the LB.

I usually share my code on GitHub, but this time I decided against it, since I haven't done anything extraordinary or special.

Results
My model gave a RMSPE of just below 0.10 on the public LB with rank 66th and RMSPE of just below 0.11 (in fact, I scored 0.10999!) which ranked me 14th on the private LB out of 3303 teams.

A lucky jump, having chosen a stable model, which results in my best individual performance on Kaggle till date, improving on my 14th rank / 2256 teams in the TFI competition.

View Public LB
View Final Results

Views
It was a tricky contest, mainly due to the nature of the public and private LB split. It was overwhelming to see so much external data being shared and used. Maybe under other circumstances, this could have played a much more important role.

Congratulations to the winner, Gert, who performed fantastically, by being way ahead of the lot in the public LB with very few submissions! And finally being stable enough to win on the private LB, again with a big lead.

So, I gained some good points from this contest, and moved to 111th in overall Kaggle rankings. My year-end goal was to be in Top-100. I'm close, and with the Walmart contest left, I might just make it.

Check out My Best Kaggle Performances

Monday, November 23, 2015

Black Friday Data Hack


AnalyticsVidhya organized a weekend hackathon called Black Friday Data Hack, which was held on 20th-22nd November, 2015.

Black Friday is actually the following weekend, but that's when we've to relax and enjoy :-)

The last hackathon was quite disappointing due to the randomness in the data and the evaluation metric. I was hoping this one would be better.
And it was. Much better.

Problem
The challenge was to predict the purchase amount of various products by users across categories given historic data of purchase amounts.

Data
In general, when you have more data, its always better. The train data had ~ 5.5 lakh observations and the test data had ~ 2.3 lakh observations. The data was very very clean and it feels wonderful to work on such datasets.

The data was of users who purchased products with the amounts. The products had data on three types of categories. The users had data about their age, gender, city, occupation, locality and marital status.

We were to build our models on the train data and score the test data which had pairs of user-product not present in the train data. The evaluation metric was RMSE, which also seemed a very appropriate choice for this problem.

Approach
I spent the first few hours just exploring the data, summarizing variables, plotting graphs, playing around with pivots and in parallel building base models (of course, XGBoost).

On the first day, I was able to go below 2500 with an optimized XGBoost model on raw features. It got me into the Top-3 and since then I've managed to maintain a position in the Top-5.

While checking the variable importance of my XGBoost, I found Product_ID was the most important variable and intuitively it made sense. So, I just submitted the average purchase amount of each product and voila! it scored 2682, which didn't seem like a very bad score. So, all those of you who couldn't cross 2682, here's a simple solution you missed.

Usually ensembles win competitions, but since I couldn't get any model close to the performance of XGB, so I decided to challenge myself to build a single powerful model. Which means, feature engineering.
These two days gave me some wonderful insights on how powerful feature engineering is. With some analysis, gut, trying, cross-validating, here are my final set of features that I used:

Model
User_ID: Used as a raw feature

User_Count: Number of observations of the user

Gender: Converted to binary

Age: Converted to numeric

Marital Status: Used as raw feature

Occupation: Used as raw feature

City Category: One-hot encoded features

Stay In Current City: Converted to numeric

Product Category 1, 2, 3: Used as raw feature

Product_Count: Number of observations of the product

Product_Mean: Average purchase amount of product

User_High: Proportion of times the user purchases products at a higher amount than the average purchase amount of the product

I built an XGBoost with these features, and the code is open-sourced on GitHub, the link is given below.

One very interesting feature I built was
F_Prop: Average purchase amount of product by female users / Average purchase amount of product by male users

This was among the top-3 important variables and gave a CV of ~ 2419 but the LB remained very similar ~ 2430, so I wasn't sure about it. I decided to go without this.

GitHub
View GitHub Repository

Results
This model gave me CV score of ~ 2425 and public LB score of 2428. I was 4th on the public LB, with Jeeban, Nalin and Sudalai in the Top-3. And we finished in the same positions with my final rank being 4th in the private LB.

Views
This is one of the best data-sets I've worked on in a while. The CV and LB scores were perfectly in sync and it was very satisfying to build features and improve the CV as well as LB scores. I'm happy with my performance as I managed to squeeze quite a lot of from the data with a single model.

I might have done better with an ensemble, but just couldn't get anything to work well. And after a while, was just too tired.

Overall, a great weekend, mostly spent on my laptop. For those of you who had memory issues, I worked on my 4GB MacBook Air throughout the weekend. Algorithms and models will advance and become optimized every day, but the power of building good features is still in the hands of Data Scientists like us.
Make the most of it until the machines come and take over ;-)

Thanks to all the folks at AnalyticsVidhya for organizing this hackathon. A big thumbs up from me.

Looking forward to the next Hackathon, and hope it gets better and more competitive.

External Links
View Other Players' Approaches on AnalyticsVidhya
View 3rd place solution code on GitHub by Sudalai Raj Kumar
View 5th place solution code on GitHub by Aayush Agrawal
 

Tuesday, October 6, 2015

World Sudoku Championship 2015


The 10th World Sudoku Championship was held on 11th-15th Oct, 2015 in Sofia, Bulgaria.

Championship Page
Download Instruction Booklet

The Indian Team was selected from the Indian Sudoku Championship and Times Sudoku Championship 2015, where Prasanna Seshadri, Rishi Puri, Kishore Kumar and I form the A-Team.

Prasanna Seshadri, Rishi Puri, Kishore Kumar and Me

"We are the best four players of the country, as we finished in the top-4 of ISC as well as TSC. It looks like a strong team with Prasanna and Rishi in good form, being finalists in the Sudoku GP and Kishore being consistent among the Indian circuit.

The Indian team stood 6th last year, which is the best performance ever, and I hope we can break into the top-5 this year. That's our goal and its probably our best shot at it in the near future, since, Rishi is not planning to continue being an active participant from next year. It will be hard to find a replacement for someone at Rishi's level, but we'll hope for the best.

On a personal level, I'm aiming for a Top-10 finish. Its been a disappointing couple of years, where I stood 14th last year and 16th the year before.
Haven't been in the best of forms lately, with some career changes and travelling going on, but hoping to give my best and make India proud!"


All of us travelled separately. Prasanna and Rishi had to reach one day before for GP Playoffs, I was flying from Mumbai, Rishi was flying from Hyderabad and Kishore was flying from Greece. Well, so much for 'best team'.

The instruction booklet just looked like a shadow of WSC 2014 in London. Very similar structure and rounds and format. I was surprised, but then realized that the main authors of WSC are Richard Stolk (Netherlands) and Yuhei Kusui (Japan) and they might've been called at the last moment to save the event.

Being a big fan of Stolk's sudokus, I was looking forward to the championship.


The rounds and points (My points vs Highest points) were as follows:

Round 1: Classics (265 vs 330)
Classics! This round went fairly smooth, I solved in reverse, attacking the hard ones first and it paid off. I scored 265 and thought that is a good start.

Round 2: Assorted (410 vs 485)
Assorted sudoku variants. That's when the aroma of Richard started. Beautiful sudokus, enjoyable round and I did well.

Round 3: Assorted (395 vs 700)
This was a bad round. I broke Inner Frame and Sum Frame, and was generally slow in a couple of other puzzles. It broke my flow and I dropped a few places.
I spent far too much time on Max Triplet (which was an excellent puzzle).

Round 4: Straight (265 vs 285)
It was expected to be a simple puzzle. It was nice, mostly got solved using row and column non-repetition. Traditionally, I've done well on such rounds (reminded me of the WSC 2012 where the Overlapping round got me into the playoffs), and I'm glad I could finish it fairly smoothly.

Round 5: Assorted (640 vs 750)
The big round! I've messed up the big round in the previous two WSCs and I really really wanted this time to be different. It was. I solved in a very nice flow, cracking one sudoku after another, without a glitch. I got a Classic wrong at the end, but still, it was a solid score, that boosted me up a few ranks.

Round 6: Assorted (375 vs 465)
The dreadful round of irregular variants. It was surprisingly good. I took the safe way, solving the easy and medium ones and leaving out the hard ones. Worked. And worked well.

So, that was the end of the Individual Rounds for Day-1. When the results came out, it was a pleasant surprise to see myself in a solid 5th position. I felt I was closing in on my dream to get into the top-5 this year.

Round 7 (Team): Relay
The team round was interesting. Sudokus were nice, and we were hoping to finish the round. Kishore and Rishi got stuck up on the Irregulars. Rishi gave up on his and Kishore didn't manage to finish his either. Prasanna had to guess on his last grid (since Rishi's Irregular relay didn't come through) which went wrong.
And to make it worse, I left two cells of Extra Region as pencilmarks, thus losing chunks of points :-(

The Great Indian Team Round Debacles continue... year after year.

Round 8: Zodiac (275 vs 625)
Ahh, feel like kicking myself. This was the only big round on Day2, and even with a mediocre performance I would've maintained my top-10 position. But it was not to be. I broke two Arrow sudokus during the round and was never able to recover from that. To make things disastrous, I swapped 6 cells in the highest pointer - Gemini, which screwed my round completely, and I fell way below 10th.

Such a disappointment. We were on track to see two Indians in the playoffs for the first time, but I messed up. Thankfully, Prasanna maintained his calm and managed to be joint 8th before playoffs.

Round 9: Multi Sudoku (140 vs 170)
I was feeling so low... and fortunately this was a low scoring round. I felt like my hands were moving in slow motion during the solve, but it wasn't too bad at the end.

So, that completed the Individual rounds of WSC 2015. I finished 14th (same rank as last year), and certainly could've done better. Maybe next year.
But that also adds to me being in Top-20 in the last 6 WSCs. Only once in Top-10 :-(

Prasanna finished 8th and made it to the playoffs, so there was something to look forward to.

Round 10 (Team): X-Killer
This was a round that most teams were looking forward to. But the organizers cancelled it due to technical issues.
Disappointing, since we had practised this well. In fact, we hosted the practise set as a contest on LMI: X-Killer
Wonderful sudokus by Deb Mohanty.

Round 11 (Team): Fractal
With our first team round going bad and second being cancelled, we had to do well on this one. It was a nice simple linked multi-sudoku, where each of us started solving from the four corners and came to the centre. We finished fairly quickly, but... The Great Indian Team Round Debacle hit us again! We swapped two digits in Kishore's corner, lost 'a lot' of points, including all the bonus.


Playoffs
Nice to see Prasanna Seshadri in the playoffs, we had an Indian there after my playoffs in 2012. He surely is a crowd-entertainer with a phenomenal performance in the playoffs first leg where he finished first among the four, thus taking him into the second leg and guaranteed 7th place.
The second leg was hard, with Bastien in 4th and having a time advantage. Bastien managed to win the leg to join Kota, Tiit and Jakub for the final leg.

Well, the same four finalists of WSC 2014 battle it out again in WSC 2015. Kota having a big lead and time advantage, raced through the playoffs, winning his second WSC crown on the trot. Tiit came in second and Jakub third, all as expected. Playoffs wasn't really exciting.
Congrats to Kota, Tiit, Jakub for the podium finishes.

Download Complete Results

Tiit Vunk (Estonia), Kota Morinishi (Japan), Jakub Ondrousek (Czech Republic)

So, Prasanna finishes 7th, which improves on my best Indian rank of 8th at the WSC. Congrats to him and this made Rishi's (so-called) last WSC memorable. Rishi finished a disappointing 38th and Kishore did well on his debut with 47th.

A-Team
7th - Prasanna Seshadri (2825)
14th - Rohan Rao (2765)
38th - Rishi Puri (1950)
47th - Kishore Kumar (1748)

B-Team and UN-Team
Rakesh Rai (1890)
Amit Sowani (1845)
Jaipal Reddy (1509)
Swaroop Guggilam (1287)
Gaurav Kumar Jain (1185)
Puneet Goenka (900)

Team India finished 9th. This is bad, considering we have been in Top-8 for the last four years. Something to dwell upon and improve.

Thanks to Richard Stolk, Yuhei Kusui, Deyan, Galya, all the other organizers and volunteers for conducting this WSC. Puzzles were fantastic, hall and seating was comfortable and overall a great experience.

Lets hope the WSC 2016 in Slovakia proves to be bigger and better. And I really hope the Indian team breaks newer and more records next year.

And guess what? India won the bid to host the World Championships in 2017! So, hoping to see you all in Bengaluru in two years time!

Saturday, September 12, 2015

Carcinogenicity Prediction of Compounds


The Carcinogenicity Prediction competition was held on CrowdAnalytix in Jul-Sep, 2015.

Objective
Carcinogenicity (an agent or exposure that increases the incidence of cancer) is one of the most crucial aspects to evaluate drug safety.

The objective was the predict the amount of carcinogenicity in compounds, which is measured through TD50 (Tumorigenic Dose rate).

Data
The train data consisted of compounds with over 500 variables consisting of physical, chemical and medical features along with their corresponding TD50 values. About 60% of the TD50 values were 0, the rest were non-zeros with few outliers.

The test data consisted of compounds with these features for which we had to predict the TD50 value.

Approach
This was a weird contest. On exploring the data, within 3-4 days, I found a key insight, and that proved to be a game changer.

So, what was this golden insight? It was the evaluation metric: RMSE.

The target variable (TD50) had many zeros and the rest were positive continuous values. RMSE as a metric can very easily get skewed due to outliers.

The train data had two values above 20,000. Predicting them accurately (greater than 20,000) would reduce the RMSE by more than 50%. So, assuming there are these outliers in the test data too, I knew this would give the maximum boost in score.

All the participants were lingering in the 1700's scores... and most of the usual models were not performing better than the benchmark 'all zeros' submission! That was a proxy validation that there had to be outliers in the test set too.

I built a model to classify outliers. The train data had only two rows (the ones with TD50 > 20,000) with target value '1' and the rest as '0'. Scored the classifier on the test set. Took the top-3 predicted rows of the test set and used 25,000 as the prediction. And BINGO! The 2nd one dropped my RMSE from 1700's to ~900. Almost a 50% drop!
Thats what you call a game-changer :-)

There are pros and cons.
Pros are that it was definitely a 'smart trick', and not really a 'sophisticated model'. Which I accepted and mentioned on the forum too. It was a neat hack applied on a poor evaluation criteria.
Cons are, of course, it doesn't lead to the best model. And worse, the result was technically determined by just one or few rows, making the rest of the test set worthless.

Model
For the remaining observations, I used a two-step model approach.

I first built a binary classifier to predict zeros vs non-zeros. Used RandomForest for this.
I then built a regressor to predict the amount of TD50, only using it for the observations which were classified as non-zeros from the binary classifier. Used RandomForest for this too.

For the binary classifier and regressor, I subsetted the train data by removing all rows where the TD50 values were > 1000 (considering them as outliers).

Results
I was 1st on the Public LB and 1st on the Private LB too.

This is my first Data Science contest where I stood 1st. Yay!
Not a really good one, but I'll take it :-)

Congrats to Sanket Janewoo and Prarthana Bhatt for 2nd and 3rd. Nice to see all Indians on the podium!

Views
The evaluation metric became the decider for this contest. A learning for me, that sometimes a simple approach can make a BIG DIFFERENCE.

Which makes it VERY IMPORTANT to explore the data, understand the objective, the evaluation and always do some sanity checks before diving deep into models and analysis. I've learnt a lot of these things from top Kagglers, and I'm sharing one of these here today, hoping someone else learns and helps in the development, improvement and future of Data Science.

Data can do magical things sometimes :-)

Check out My Best CrowdAnalytix Performances

Saturday, September 5, 2015

Puzzle Ramayan 2016

The online rounds of Puzzle Ramayan 2015-2016 have ended! This is a national level event aimed at encouraging puzzle solvers of India to participate and compete with the top solvers to gain experience and improve competition in the years to come.

NOTE: This event serves as a qualifier to participate in the Indian Puzzle Championship 2016

The championship consisted of 8 online rounds (Sep-2015 to Mar-2016) from which the top solvers will be invited to participated in the national finals.
Championship Page

National Finals
The finals will be held on 17th July, 2016 in Chennai.

View Finals Page


Online Top-10
1. Rohan Rao - 597.3
2. Amit Sowani - 575.8
3. Swaroop Guggilam - 477.4
4. Rajesh Kumar - 458.3
5. Rakesh Rai - 411.2
6. Ashish Kumar - 374.2
7. Kishore Kumar - 372.1
8. Jayant Ameta - 344.2
9. Jaipal Reddy - 302.7
10. Devarajan D - 277.2

View Complete Results

P.S. Prasanna's name is removed from list since he has a wild card for the WPC next year on being the best Indian performer at the WPC this year.


Round 8: Placement (26th - 28th Mar, 2016)
Author: Rajesh Kumar
Download Instruction Booklet
Download Puzzle Booklet

View Results
View Forum

Nice puzzles, but on the harder side. A little disappointed that I wasn't able to finish the set.
Horrible answer keys, I struggled a lot with it, so did a few other players.

Overall, a decent end to PR.

The Top-10 look more-or-less as expected, but really good to see Ashish and Kishore improving and a great job by Devarajan for maintaining his top-10 position throughout the rounds. Looking forward for an interesting and fun-filled finals in Chennai in July.


Round 7: Loops (27th - 29th Feb, 2016)
Author: Prasanna Seshadri
Download Instruction Booklet
Download Puzzle Booklet

View Results
View Forum

Wow! Another wonderful round. I'm not very good at Loops, but I could solve this very smoothly. Puzzles were excellent, and a very well-balanced set and PR round. Probably the best so far.

I finished the set in 47mins, with Swaroop in 57mins and Amit in 65mins.  Swaroop now has increased his lead at 3rd place above Rajesh and should be able to hold on to it since the last round is authored by Rajesh.

Hope to end it well.


Round 6: Shading (23rd - 26th Jan, 2016)
Author: Swaroop Guggilam
Download Instruction Booklet
Download Puzzle Booklet

View Results
View Forum

Wonderful! What a perfectly balanced round this was. Kudos to Swaroop for authoring this set, my favourite round of PR so far. Prasanna finished the set in 44mins, I finished in 58mins and Amit in 65mins.

Lot of swaps in the points table after this round. Also due to the rankings being updated after discarding the worst two scores. I regain the top spot over Amit. Rajesh is less than a point above Swaroop. A disappointing round by Rakesh allowed Kishore to move above him.

With the last two rounds to go, it will be an interesting finish, especially crucial for Swaroop, who needs to be in the Top-3 to be eligible for the NRI wildcard.


Round 5: Snake (26th - 28th Dec, 2015)
Author: Ashish Kumar
Download Instruction Booklet
Download Puzzle Booklet

View Results
View Forum

After a great Round 4, I had an absolutely disastrous Round 5. Snakes is not a type I really enjoy, and it showed here. I scored a poor 57 points, compared to Amit's 88. Prasanna did well by finishing all puzzles just within 90 minutes.

Puzzles were top-notch quality from Ashish, but they were too hard for PR. I'm not surprised to see the participation low, but a little surprised by some regular names missing, including ones in the current Top-10.

Lot of changes in the top-10 after this round. Amit takes the top spot with a good lead, Swaroop moves to 3rd, above Rajesh, and finally, Rakesh moves above Kishore.
Its getting interesting, and I hope the remaining 3 rounds are better, way better. 


Round 4: Regions (28th - 30th Nov, 2015)
Author: Rakesh Rai
Download Instruction Booklet
Download Puzzle Booklet

View Results
View Forum

This was one of my better performances in a puzzle contest in recent times. The puzzles were of my liking. Yin Yang, Spiral Galaxies, Area Division are in my all-time favourites, and it was wonderful to solve this set. Puzzles were really fun and it was a better set than the last 3 PR rounds.

I topped the round by finishing in 48mins and was ranked 11th internationally, which is my best rank after Twist way back in 2011. Amit did well by finishing in 56mins, Prasanna finished in 64mins and Swaroop in 83mins.

That put me on top in PR rankings and also got me my best LMI Rating in Puzzles! So, a pretty good weekend!


Round 3: Evergreens (31st Oct - 2nd Nov, 2015)
Author: Amit Sowani
Download Instruction Booklet
Download Puzzle Booklet

View Results
View Forum

That was hard! Especially for the type of rounds expected in PR. Well, it was time to improvise. Since this resulted in a very low scoring round, the scoring system was changed to add a bit of normalization so that such variability in the difficulty of tests can be overcome to some extent.

Even though I topped the round (among Indians), it didn't feel like a smooth performance. Felt like I could've added some 8-10 points more to my score of 73.

Prasanna tested the puzzles, so you won't his name on the scorepage. Congrats to Rajesh and Swaroop for their good performances.


Round 2: Number Placement (26th - 28th Sep, 2015)
Author: Deb Mohanty
Download Instruction Booklet
Download Puzzle Booklet

View Results
View Forum

I didn't do too well. Couldn't finish the round, and got stuck up in too many puzzles during the test.
Puzzles were really nice. Much better than Deb's SM round :-)

Congrats to Prasanna who finished the set in 72 minutes and Amit who just managed to finish it before time. I scored 97.4 points.

This was supposed to be one of the rounds I was most comfortable with, and it bombed. Hope to cover-up in the next few rounds. I also hope this is the worst performance which will get discarded (along with R1 which I authored).


Round 1: Classics (5th - 7th Sep, 2015)
Author: Rohan Rao
Download Instruction Booklet
Download Puzzle Booklet

View Results
View Forum

Congrats to Prasanna, Swaroop and Amit for completing the set. Prasanna finished in 50mins which put him in 12th place worldwide. Swaroop and Amit were very close and finished just one second apart in the 77th minute.

Its nice to see my three team-mates, who will represent India at the WPC along with me, performing at the top among the Indians.

Congrats to Endo, Ulrich and Hideaki who take the top-3 international spots.

Overall, I'm glad the feedback was positive and most participants enjoyed the puzzles. There was some discussion around one puzzle, Hitori Blocks, being a tad harder than the rest for this set. I agree it was a little outlier, but it didn't affect rankings and performances much. Most of the results were as expected.

Seems like a good start to PR... 57 Indians with non-zero scores and totally 304 participants. I hope these numbers increase in subsequent rounds. And I'll be participating in the coming rounds! :-)