Kaggriculture: Kaggle Farming Agent
A bot that runs a farm against an opponent in a shared market, built for the Kaggle × Google Kaggriculture simulation competition. Reached rank 423 of 9,984 (top 5%).
Overview
Kaggriculture was a seven-week Kaggle × Google simulation competition with a $50,000 prize pool and about 10,000 teams. I built two agent architectures and an imitation-learning pipeline, and my own agent peaked around the top 40%. After finding that my local engine no longer matched the live rules, I adopted prvsiyan's open-source bot (Apache-2.0, with attribution) and spent the final week improving it. My upgrade won 29 of 34 fresh ladder replays where the original won 18, and took the team to rank 423 of 9,984 with a score of 2,513.7.
The problem
Each game is 720 one-second turns: buy seeds, animals, land and farmhands, then grow, harvest and sell to finish with the most cash. Both players sell into one shared market, so every sale lowers the price for both of them. That makes it an economic game as much as a production problem, ranked over thousands of bots playing each other around the clock.
Key features
- Route planner that harvests and sells in the same afternoon
- Crop valuation at the price on the day each crop reaches market
- Exact per-unit market accounting and a replayer accurate to the dollar
- Imitation-learning pipeline trained on 732 expert games
- Publish gates that block changes that fail testing
What I did
- Built two agent architectures from scratch (an economic planner, then a route planner worth about +170 rating)
- Trained a CNN on 4.8 million expert decisions and kept it out of the agent when it made results worse
- Found that my local engine no longer matched the live rules by replaying real ladder games
- Upgraded the strongest open-source bot, lifting fresh-replay wins from 18 to 29 of 34
- Ran a final week of parallel AI-agent investigations, each checked by a skeptic agent
AI under the hood
I built this with AI coding agents, mainly Claude Code, with Gemini as an outside reviewer and Muse Spark building and submitting one agent end to end. That meant an imitation-learning pipeline in two days, 150+ candidate agents, and a final week of seven parallel investigations, each checked by a skeptic agent, while I made the calls.
The failure mode was over-claiming from thin evidence. What kept it honest was process: a running log of retractions, negative results written next to the ideas they killed, skeptic agents whose job was to refute, and publish gates that could say "don't ship this".