Training a search agent is great for RL because every lever impact model behavior in legible ways. Jasper's guide shows how small updates to the reward function teach a model to avoid sloppy tool calls, prune unnecessary docs, and balance persistence with token efficiency.
Jasper's guide shows how reward tweaks shape search agent behavior
AISummary
Jasper Lu's new blog post walks through training a search agent with GRPO, showing how small reward function changes teach a model to avoid sloppy tool calls, prune unnecessary documents, and balance persistence against token efficiency.
The post makes every rollout browsable and releases the code as open source, with the full process from learning rate sweeps to reward shaping documented.
Post on XView on X
@tinkerapi
Sharing my first of hopefully many research blog posts! https://jasperlu.com/blog/training-search-agents-grpo/ This one is the kind of educational blog post I wish I'd had when I started with RL for LLMs. I tried to make it as open as possible. Every rollout is browsable, the code is open source, and I walk through my entire thought process, from learning rate sweeps to reward shaping.View quoted post on X
Source: Tinker · x.comPublished · added here
