Blog

Notes on post-training, self-improvement, and agents.

In preparation

Notes I am currently preparing:

Exploration in the LLM Space

A look at exploration in classical reinforcement learning, why it becomes a bottleneck in reinforcement learning for language models, and which ideas may be worth trying next.

Memory vs. Parameters

An analogy to how humans learn, followed by a recap of research on updating a model through its parameters or through external memory.

Anthropic's Full Ecosystem

A researcher's tour of the Claude ecosystem: models, Claude Code, the Agent SDK, MCP, and skills, and what it signals about where AI systems are heading.

Want a note when the first one lands, or want to discuss any of these topics before I write them? Email me; I'd genuinely enjoy that.