- Before a new Claude model ships, a small group of customers is already testing it, breaking it, and shaping what ships with it — and this short film captures what those teams are learning.
- They describe the energy of early access as "all hands on deck," moving at the speed of light, and treating it as a generational opportunity that comes with a sense of responsibility.
- The first things they throw at a new model are automated evals running in the background and hard real-world tasks (like drafting an S1), and seeing previously-failing evals suddenly start working consistently is the sign a model is going to be special.
Before we ship a Claude model, these teams try to break it.
- A small group of customers tests and tries to break each new Claude model before launch, helping shape what ships with it.
- The energy of getting something new from Anthropic is described as exciting like an approaching storm — all hands on deck, moving at the speed of light, dropping whatever you're working on.
- Working at the frontier feels like being in constant learning mode, a generational opportunity that brings both luck and a sense of responsibility to push the envelope, innovate, be more secure, and make things easier to build with.
- The very first thing teams do with a new model is start automated evals running in the background.
- A pipe-dream complex legal task cited is drafting an S1; with agentic capabilities (finding information, synthesizing it, editing documents), models can now handle larger and larger chunks of it autonomously.
- Just by swapping in one new model, an agent went from sometimes answering and sometimes getting stuck to answering every question quickly and accurately — and a testing agent's success-rate dashboard jumped about 20%.
- Tasks that don't work today are the best sign of what the next models will be far better at; seeing evals that never worked start working consistently signals a model is going to be something special.
- The relationship with Anthropic feels collaborative ("we build with you," near-daily conversations, engineers like they're on the same team) with a high trust bar that anything published won't be "slop"; one word to characterize building at the frontier was "dazzling" — bright, full of opportunity, with bigger waves coming.
Frontier — The leading edge of AI capability where the newest, most capable models are developed and tested.
Eval — An evaluation (often automated) that measures how well a model performs on specific tasks.
Agentic capabilities — A model's ability to autonomously find information, synthesize it, and edit documents to complete tasks.
S1 — A complex registration document filed with regulators (e.g. for an IPO), cited as a hard legal drafting task.
Testing agent — An agent whose success rate is tracked on a dashboard to measure model performance gains.
Slop — Low-quality output; the customers hold a high trust bar that what Anthropic ships is not slop.
Compounding — The reinforcing cycle where the latest tools improve customers' products, which improves their products in turn.
Before a new Claude model ships, a small group of customers is already testing it, breaking it, and shaping what ships with it. We sat down to see what they're learning.
When you get something new from Anthropic, what is the energy like? We know a storm's ahead, but there's something exciting about a storm because it's all hands on deck. It feels like we're moving at the speed of light. It's like getting the call and jumping from whatever you're working on. We have something new, let's figure out what it's like. The moment we get a new model from Anthropic, we realize the grounding has changed.
What's it like to work at a company that's helping to shape the frontier? It's insanely fun. All of us are just in learning mode. This moment just feels like a generational opportunity for anyone in this industry. I feel very lucky and also very responsible. We need to continue to push the envelope, continue innovating, being more secure, and making things easier to build with. In a way, I love that I can unlock a new class of developers and builders.
What's the first thing you throw at a new model? The very first thing is we will start automated evals just so that they start running in the background. One use case that is a pipe dream that's easy to point to as a particularly complex legal task is drafting an S1. Now with agentic capabilities where these models can go out and find information that they need, synthesize it, edit documents, we're getting to larger and larger chunks of the S1 that you can just send the model on its way to do. Just by swapping in that one model, every question I ever wanted to ask it started getting answered. It went from this agent can sometimes answer questions, sometimes get stuck, to, oh my God, it is answering every question quickly and accurately. The dashboard of the testing agent success rate has just increased by, I think it's 20%.
Things that don't work today are the best sign for, here's what the next models are going to be way better at. Seeing evals that have never worked start working and then start working consistently, this model is going to be something special.
What's it like working with Anthropic? It feels like I have a conversation with you almost every other day. The engineers on the team, I feel like, are almost on the same team. It's less like we're just buying something from you, and more like we build with you. We have a very high trust bar that anything you publish is not going to be slop.
What is one word or phrase that characterizes what it feels like to actually be building at the frontier? Dazzling, if that makes sense. It can be blinding at times. Just the brightness, opportunity, excitement. Compounding — we get the latest tools, which leads to our customers getting a better product, which leads to us getting better products. You have a big wave under you that is changing the way your user is working and changing the way you are working. And you have to keep your balance. And you know there are bigger waves coming.
TL;DR
- Before a new Claude model ships, a small group of customers is already testing it, breaking it, and shaping what ships with it — and this short film captures what those teams are learning.
- They describe the energy of early access as "all hands on deck," moving at the speed of light, and treating it as a generational opportunity that comes with a sense of responsibility.
- The first things they throw at a new model are automated evals running in the background and hard real-world tasks (like drafting an S1), and seeing previously-failing evals suddenly start working consistently is the sign a model is going to be special.
Takeaways
- A small group of customers tests and tries to break each new Claude model before launch, helping shape what ships with it.
- The energy of getting something new from Anthropic is described as exciting like an approaching storm — all hands on deck, moving at the speed of light, dropping whatever you're working on.
- Working at the frontier feels like being in constant learning mode, a generational opportunity that brings both luck and a sense of responsibility to push the envelope, innovate, be more secure, and make things easier to build with.
- The very first thing teams do with a new model is start automated evals running in the background.
- A pipe-dream complex legal task cited is drafting an S1; with agentic capabilities (finding information, synthesizing it, editing documents), models can now handle larger and larger chunks of it autonomously.
- Just by swapping in one new model, an agent went from sometimes answering and sometimes getting stuck to answering every question quickly and accurately — and a testing agent's success-rate dashboard jumped about 20%.
- Tasks that don't work today are the best sign of what the next models will be far better at; seeing evals that never worked start working consistently signals a model is going to be something special.
- The relationship with Anthropic feels collaborative ("we build with you," near-daily conversations, engineers like they're on the same team) with a high trust bar that anything published won't be "slop"; one word to characterize building at the frontier was "dazzling" — bright, full of opportunity, with bigger waves coming.
Vocabulary
Frontier — The leading edge of AI capability where the newest, most capable models are developed and tested.
Eval — An evaluation (often automated) that measures how well a model performs on specific tasks.
Agentic capabilities — A model's ability to autonomously find information, synthesize it, and edit documents to complete tasks.
S1 — A complex registration document filed with regulators (e.g. for an IPO), cited as a hard legal drafting task.
Testing agent — An agent whose success rate is tracked on a dashboard to measure model performance gains.
Slop — Low-quality output; the customers hold a high trust bar that what Anthropic ships is not slop.
Compounding — The reinforcing cycle where the latest tools improve customers' products, which improves their products in turn.
Transcript
Before a new Claude model ships, a small group of customers is already testing it, breaking it, and shaping what ships with it. We sat down to see what they're learning.
When you get something new from Anthropic, what is the energy like? We know a storm's ahead, but there's something exciting about a storm because it's all hands on deck. It feels like we're moving at the speed of light. It's like getting the call and jumping from whatever you're working on. We have something new, let's figure out what it's like. The moment we get a new model from Anthropic, we realize the grounding has changed.
What's it like to work at a company that's helping to shape the frontier? It's insanely fun. All of us are just in learning mode. This moment just feels like a generational opportunity for anyone in this industry. I feel very lucky and also very responsible. We need to continue to push the envelope, continue innovating, being more secure, and making things easier to build with. In a way, I love that I can unlock a new class of developers and builders.
What's the first thing you throw at a new model? The very first thing is we will start automated evals just so that they start running in the background. One use case that is a pipe dream that's easy to point to as a particularly complex legal task is drafting an S1. Now with agentic capabilities where these models can go out and find information that they need, synthesize it, edit documents, we're getting to larger and larger chunks of the S1 that you can just send the model on its way to do. Just by swapping in that one model, every question I ever wanted to ask it started getting answered. It went from this agent can sometimes answer questions, sometimes get stuck, to, oh my God, it is answering every question quickly and accurately. The dashboard of the testing agent success rate has just increased by, I think it's 20%.
Things that don't work today are the best sign for, here's what the next models are going to be way better at. Seeing evals that have never worked start working and then start working consistently, this model is going to be something special.
What's it like working with Anthropic? It feels like I have a conversation with you almost every other day. The engineers on the team, I feel like, are almost on the same team. It's less like we're just buying something from you, and more like we build with you. We have a very high trust bar that anything you publish is not going to be slop.
What is one word or phrase that characterizes what it feels like to actually be building at the frontier? Dazzling, if that makes sense. It can be blinding at times. Just the brightness, opportunity, excitement. Compounding — we get the latest tools, which leads to our customers getting a better product, which leads to us getting better products. You have a big wave under you that is changing the way your user is working and changing the way you are working. And you have to keep your balance. And you know there are bigger waves coming.