Rohan LopesRohan LopesSenior Integration & AI Architect
HomeWorkBlogAboutContact

Rohan Lopes

Senior Integration & AI Architect. Building integrations and AI agents that do real work.

HomeWorkBlogAboutContactBlog archive

© 2026 Rohan Lopes · Built with Next.js and Claude Code · Privacy

    All posts
    Kaggle
    Game AI
    Agents
    Simulation
    Imitation Learning
    Python
    PyTorch
    Claude Code
    Muse Spark

    Kaggriculture: Seven Weeks, One Farming Bot, and a Lesson in Measurement

    A retrospective on the Kaggle × Google Kaggriculture competition: two agent architectures, an imitation-learning pipeline, and a final week of AI-assisted analysis. The biggest lessons were about measurement, not farming.

    October 1, 2026 · 8 min read

    TL;DR. For seven weeks I competed in Kaggriculture, a Kaggle × Google simulation competition with a $50,000 prize pool. You write a bot that runs a farm against an opponent in a shared market. I built two agent architectures (an economic planner, then a route planner) and an imitation-learning pipeline, and my own agent peaked around the top 40%. In the final week I took the best open-source bot and upgraded it: my latest submission isn't the notebook as-is, but an improved version that won 29 of 34 fresh ladder replays where the original won 18. On the live ladder it quickly overtook the unmodified copy I kept as a safety net, and took the team to rank 423 of ~10,000, with a score of 2,513.7. Most of what I learned was about measuring things correctly: my simulator, my tools and my test opponents were each wrong in ways that mattered more than any strategy decision.

    At a glance

    CompetitionKaggriculture: Kaggle × Google, $50,000, ~10,000 teams; I competed 5 Aug–24 Sep 2026
    RoleSolo entrant, working with AI coding agents
    Final submissionMy upgraded version of prvsiyan's open-source bot, with an unmodified copy kept as a safety net
    ResultRank 423 of 9,984 (top 5%), score 2,513.7, as of 25 Sep (leader: 3,112.4); final standings on the leaderboard
    Own agentBuilt from scratch; peaked around the top 40%
    Volume32 submissions · 150+ candidate agents · ~1,000 real ladder games analysed · 732 expert games for imitation learning
    StackPython, NumPy, PyTorch, kaggle-environments, Kaggle API, Claude Code, Gemini, Muse Spark

    The game

    Each game lasts 30 in-game days of 24 hours: 720 turns, one second each. You start with $3,000, buy seeds, animals, land and farmhands, and grow, harvest and sell. Score is cash at the end. The twist is that both players sell into one shared market: every unit you sell lowers the price for both of you, while town shops slowly buy stock back. It's a production problem and an economic game at once, ranked over thousands of bots playing each other around the clock.

    Where it ended

    In mid-September I found that my local game engine no longer matched the live one (Act 6 below). Every strong public notebook beat my best agent by about $40,000 a game, so I made a pragmatic call: I adopted the strongest open-source bot, prvsiyan's "frontier" (notebook, Apache-2.0, submitted with attribution), and spent the last week improving it.

    That week I used AI agents at scale: seven parallel investigations, each checked by an independent skeptic, then a critic hunting for gaps. They covered our own ladder games, the top-10 teams' play, new public notebooks, the engine's rules and the ranking system. The conclusions were sobering and useful:

    • The top 10 run private bots that react to their opponent. They can't be copied from replays, and they win three games in four or more against our rating band.
    • Most of our games were against close relatives of our own bot, and in our current band 70% of decided games were within $1,000. Small, well-aimed edges flip real games.

    My latest submission is therefore not the open-source bot as-is but my upgraded version of it, which adds a handful of small changes. Most help it sell before rivals running the same plan, or stop it losing stock to the shed limit. Two, a sales forecast and a feeding safeguard, are ported from arsgorynich's open herd-safe-v3 notebook. I tested it two ways: on real ladder games played after tuning ended, replayed with the new bot in our seat, and against public bots that react normally. In the replays, private opponents repeat their recorded moves and can't respond, which flatters a sell-timing change.

    TestFrontier (before)Final bot
    34 fresh ladder games, never used for tuning18 wins, 13 losses, 3 ties29 wins, 5 losses
    12 public bots, 144 local games88% wins95% wins

    The live ladder agreed. After about 100 games the upgrade was rated 2,513.7, against 2,363 for the unmodified copy, and it moved the team from around rank 670 to 423 of 9,984. It won't reach the top 10 (around 2,944), which needs a different class of bot, and much of the rating still belongs to the open-source authors whose work I built on. Both bots stay live until the ladder settles, and Kaggle ranks the team by the better of the two.

    What worked, and what didn't

    WorkedDidn't
    Reading the engine source instead of the rulesImitation learning as a policy
    Valuing crops at their sale-date priceCopying elite boards piece by piece
    Last-day liquidation and respecting shed capacityFront-running and unilateral restraint
    Rebuilding around a route planner (+170 rating)A macro-planner and four sell schedulers
    Exact per-unit market accountingInstalling the "optimal" sell schedule
    Publish gates that could say noTuning against self-play and recorded tapes
    Replaying our real games to test changesTrusting early ladder ratings

    Lessons I'm keeping

    1. Validate the simulator against real games, early and again later. The live rules changed under me, and a routine replay check would have saved weeks.
    2. Audit your instruments. My own tooling had bugs that quietly corrupted results for weeks. Tools that label your own data should fail loudly, not guess.
    3. Your test opponents define what you optimise. Self-play rewards beating yourself; recorded opponents can't react. Test against the field you'll actually face.
    4. An accounting gap isn't reachable money. A $44,000 "sell-timing prize" shrank to $8,000 once storage limits were respected, and installing it lost money.
    5. Partial imitation is worse than either coherent strategy. A winning board is the output of an economy; copying its surface without the engine makes you worse.
    6. Shared markets are games, not optimisation problems. Some moves help both players; unilateral restraint gets punished.

    Working with AI

    I built this with AI coding agents, mainly Claude Code, with Gemini as an outside reviewer and Muse Spark building and submitting one agent end to end. The speed was real: an imitation-learning pipeline in two days, 150+ candidate agents, and a final week of seven parallel investigations, each checked by a skeptic agent, while I made the calls.

    The failure mode was just as real, and it matched my own: over-claiming from thin evidence. What kept it honest was process: a running log of retractions, negative results written next to the ideas they killed, skeptic agents whose job was to refute, and gates that could say "don't ship this". The best results of the project came from that loop, not from any single clever idea.


    The full story

    Act 1: read the source, not the rules

    The engine shipped as Python, and the prose rules left out most of what decides games. Melon is a one-shot ~$26,000 pool no shop restocks, so whoever sells first wins it. The shed holds 100 units and destroys the overflow at midnight. A planner that valued each crop at the price it would fetch on the day it reaches market won 28 of 28 games against my hand-tuned build.

    Act 2: imitation learning, 94.6% accurate and still worse

    Top players' best games reached ~$180,000 to my ~$60,000, so I trained a small CNN on 732 of their games (4.8 million decisions). It predicted their moves 94.6% of the time, and the agent got steadily worse the more it trusted the model: from $75,000 with the network off to $36,000 when it made most decisions. The experts' choices assumed their farm, not mine. The publish gate never said yes, which is exactly what it was for.

    Act 3: grinding the ladder

    August was a loop of hypothesis, local test, submit and wait. Some changes helped: last-day liquidation, holding three land quadrants instead of four, and not counting on town shops that hadn't opened yet. Others didn't, and my self-play harness predicted the ladder backwards twice.

    Act 4: the top of the ladder looked like a recording

    Two top players' games were identical for 697 of 720 turns: in August, much of the ladder above me ran recorded action tapes from public notebooks, which works because every game starts from the same board (by September the top 10 had moved on to reactive private bots). I rebuilt my agent as a route planner that harvests and sells in the same afternoon. It beat my old code by about 170 rating points side by side, the biggest gain my own code ever made.

    Act 5: exact tools, and the $44k that was really $8k

    I built exact market accounting and a replayer accurate to the dollar. A dynamic-programming bound promised $44,000 a game from perfect sell timing, but the optimum stored far more than the shed can hold. The realistic prize was about $8,000, and installing the schedule wholesale cost $72,000. By mid-September nothing I tried beat my baseline.

    Act 6: the engine was wrong

    Then I replayed a real ladder game through the current official engine. It matched to the dollar, while the engine every experiment had run on did not: the live rules had quietly changed. On the correct engine my agent lost every game to the public notebooks, by about $40,000 each.


    The standings will keep moving until the ladder settles in mid-October; follow them on the Kaggriculture leaderboard (team "Rohan L").

    Thanks to Kaggle and Google for a genuinely deep environment, and to the authors of the public notebooks, especially prvsiyan, tetsutani and arsgorynich, whose open work taught me more in a week than my own did in a month.

    Project

    Kaggriculture: Kaggle Farming Agent

    All posts