Exploration in the LLM Space
A look at exploration in classical reinforcement learning, why it becomes a bottleneck in reinforcement learning for language models, and which ideas may be worth trying next.
Notes on post-training, self-improvement, and agents.
Notes I am currently preparing:
A look at exploration in classical reinforcement learning, why it becomes a bottleneck in reinforcement learning for language models, and which ideas may be worth trying next.
An analogy to how humans learn, followed by a recap of research on updating a model through its parameters or through external memory.
A researcher's tour of the Claude ecosystem: models, Claude Code, the Agent SDK, MCP, and skills, and what it signals about where AI systems are heading.
Want a note when the first one lands, or want to discuss any of these topics before I write them? Email me; I'd genuinely enjoy that.