- The Code with Claude Tokyo 2026 keynote announces Anthropic's fifth-generation models, Claude Fable 5 and Claude Mythos 5, their most capable models ever, and frames Anthropic as a platform company that helps developers close the gap between exponential model capability and linear business adoption.
- Speakers cover three layers — models (Diane), the Cloud Platform and Cloud Managed Agents (Angela and Caitlyn), and Claude Code (Cat) — with new managed-agent features (scheduled deployments, environment-variable vaults, memory, dreaming) and Claude Code surfaces (desktop app, agents view, dynamic workflows).
- Customer stories (Rakuten, Canva, Notion, Asana, Spotify, Mercari, Williams F1) illustrate AI-native companies where work runs on AI and people decide outcomes, and developers are urged to design for the next version of Claude, not just today's.
Code with Claude Tokyo 2026: Opening Keynote
- Anthropic released its fifth-generation models, Claude Mythos 5 and Claude Fable 5; Fable 5 is the most capable generally available model and is state-of-the-art on nearly all tested AI benchmarks, while Mythos 5 is the same model with cyber and bio safeguards lifted, available to Project Glasswing partners.
- Anthropic positions itself as a platform company: year-over-year API volume is up nearly 17x, and the platform gives developers tools to build agents because most people will experience AI through products others build on it.
- Model capability is on an exponential while most business capability is linear, creating a growing gap that the Claude platform aims to close; the keynote traces progress from drafting commit messages to Mythos finding a 27-year-old OpenBSD vulnerability.
- Fable 5 excels at long-horizon autonomy (running for days on one goal across millions of tokens, dispatching sub agents), single-shot correctness, reading code, vision, and non-coding knowledge work; a new safeguard routes cyber/bio/chem requests to Opus 4.8.
- Cloud Managed Agents combine three ingredients — an agentic harness (separating the "brain" from sandbox "hands"), context (1M-token window, memory, self-written skills, "dreaming"), and production infrastructure (auto-scaling sandboxes, agentic fleets).
- New managed-agent features shipping today include scheduled deployments (run agents on any cadence) and environment-variable vaults for secure authenticated API requests, demonstrated on a fictional Shankiro racing dashboard with outcomes-based rubrics.
- Claude Code now spans the CLI, IDE, the Cloud Desktop app, and a new agents view, supporting "multi-clauding" in parallel; Anthropic engineers ship 8x more code than in past years, and products include code review, remote control on iOS/Android, routines, Claude security, and dynamic workflows.
- A dynamic-workflows demo localizes a marketing site into 12 languages in parallel (then verifies each), versus nearly an hour sequentially; customers like Spotify (1,000+ PRs/month, migration time cut 90%) and Mercari (engineering output up 90% YoY) show org-scale adoption.
Claude Fable 5 — Anthropic's most capable generally available fifth-generation model, state-of-the-art on nearly all tested benchmarks.
Claude Mythos 5 — The same underlying model as Fable 5 with cyber and bio safeguards lifted, available to Project Glasswing partners.
Cloud Managed Agents — Anthropic's product for building and deploying long-running agents with a harness, context tools, and production infrastructure.
Harness — The tools, environment, and permissions that let a model actually do work, separating the deciding "brain" from the executing sandbox "hands."
Dreaming — A managed-agent feature where an agent reviews past sessions to update its memory and skills and self-improve.
Project Glasswing — Anthropic's program for safely releasing high-capability models, with a safeguard system routing sensitive requests.
Dynamic workflows — A Claude Code feature that runs Claude in parallel across many agents in a deterministic structure for large-scale tasks.
Time horizon — How long a model can work autonomously before losing coherence on what to do next.
AI-native company — A company where work itself runs on AI and people decide what the outcomes should be.
Please welcome to the stage head of engineering for the Claude platform at Anthropic, Caitlyn Les. Tokyo, this is the first time that we've brought Code with Claude to Japan and we're grateful to spend the next couple of days with all of you. This morning we'll be talking about our models, our platform, and our products. But before we get into it, I want to start by sharing that just a few hours ago, we released the fifth generation of Claude models, Claude Mythos 5 and Claude Fable 5. These are our two most capable models ever. Diane, our head of product for research, will join us shortly to share a lot more about why these models are so special.
I lead engineering for the Claude platform. The platform gives developers the tools they need to build systems on top of Claude to harness its intelligence. This is the highest-leverage way for us to help solve the world's most important problems. And this is why Anthropic is a platform company. Developers all over the world, many in this room today, produce far more value on top of the platform than we could ever build on our own. So let's start with what I'm seeing from our customers lately. There's an incredible volume of powerful applications shipping right now, and a lot of it is coming from right here in Japan. Rakuten is one of our favorite customers to work with. Their team went from using Claude Code to accelerate development to building on Cloud Managed Agents to power custom internal agents across engineering, product, sales, and finance. One of their product managers coordinates teams of agents exactly the same way a leader manages teams of humans. Now they're shipping major releases every two weeks instead of only once per quarter. And it's the same across Asia-Pacific. Another great customer is Canva, the Australian design platform that hundreds of millions of people use. Most of their users have never written a line of code. But with the help of Claude, Canva Code changes all of this. Within a design, you can just ask for something interactive like a map, a calculator, or a widget, and Claude builds it into a working mini app ready to drop right into your page.
All around the world, people are building new systems and applications powered by Claude that nobody could build before. The landscape is changing faster than ever. Things that weren't possible yesterday have become possible today. We're on a mission to keep raising the ceiling of what's possible to build by making models that are increasingly capable. A couple of years ago, the frontier of model development was just the ability to draft a really simple commit message. One year ago, we were standing on stage at our first-ever Code with Claude event. Opus 4 was the headline, and it was mind-blowing at the time that Claude could build an entire feature on its own. Six months ago, agents were able to run overnight to complete long-running and autonomous tasks. Two months ago, Mythos read the entire OpenBSD source tree and found a 27-year-old vulnerability that had slipped past human reviews and static analysis for almost three decades. And earlier today, we released Mythos 5 and Fable 5. Fable 5 is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, scientific research, vision, and more. These jumps keep getting bigger, but the time intervals keep getting shorter. Even though these model capabilities are improving on this exponential, most business capabilities are still on a linear. So there's a growing gap between what AI can do and what it's actually doing for people. Closing that gap is our collective opportunity, and that is why we've built the Claude platform. Year-over-year API volume is up nearly 17 times on the platform.
This morning, Diane will talk about our models. Next, Angela and I will walk you through how you can build and deploy agents at scale on the Cloud Platform using Cloud Managed Agents. Today we're excited to share that we're shipping brand new Cloud Managed Agents features. You can now schedule deployments to have your agents run on whatever cadence you need, and you can store environment variables in vaults so your agents can securely make authenticated API requests without giving them access to the keys. Finally, Cat will walk you through the latest in Claude Code, where the average developer is now spending 20 hours per week with Claude. Most people will never integrate with the Claude API themselves. They'll experience AI through something that one of you built on the Claude platform, like a salesperson walking into a high-stakes meeting fully briefed by Slack agents, or a lawyer getting a brief out the door faster than ever with Legora. This is why we're a platform company. Thank you for being here, for partnering with us, and for showing us what's possible. Arigato gozaimasu.
Please welcome Diane from our research product team. Hi, good morning Tokyo. I'm Diane and I joined Anthropic in 2023, and I've been a part of every version of Claude since Claude 2. That's bringing 21 versions of Claude across Haiku, Sonnet, Opus, and now Fable and Mythos to end users and developers like you. Our most recent launch happened just a few hours ago. We released Claude Fable 5 and Claude Mythos 5, the first generation of our fifth models. Fable 5 is the most capable model we've ever made generally available. It's based on the same foundations as Mythos 5. We talk a lot about the exponential at Anthropic. As model intelligence increases, we believe the value of use cases it creates increases exponentially. For example, agentic coding that many of you are using today is far more valuable than simple autocomplete from just a few years ago.
Let's start with coding, because we believe that's where most of you will feel the magic first. Fable 5 performs the highest on SWE-bench Pro. But the benchmark itself undersells the story. The longer and more complicated and sophisticated the task, the farther the gap between Fable and every other model out there. Two things drive that gap. The first is single-shot correctness. Give Fable 5 a complex, well-specified problem and it will nail it on the first pass. Early testers have told us about single prompts that essentially create work that would have taken days or even weeks for a group of teams. The second is long-horizon autonomy. Fable 5 can run for days on a single goal and stay coherent the entire way. It'll remember your specifications even on tasks that span millions of tokens. And it can dispatch sub agents, keep them on track far more dependably and with more cost consciousness than any other model we've shipped. One more thing: Fable 5 isn't just good at writing code, it's even better at reading it. It's better at triaging outages and digging through your repo history to figure out what's broken when, and proactively surfacing suggestions for making it better. Fable 5 is just as much of a step outside of coding. Whether that's financial analysis, documents, slides, spreadsheets, Fable 5 manages that work end to end. It'll follow instructions, stay on scope, and what it creates back for you will be professional grade. It's also better where requests aren't clean. Fable 5's vision is also the best in the industry. It can read dense technical images, web applications, plots, diagrams, and charts far more accurately than any other version of Claude.
Now, intelligence of this magnitude cuts both ways. Two months ago, we launched Project Glasswing and made Mythos preview available to a small group of partners because its capabilities in cybersecurity were strong enough to be potentially misused. Since then, we built a new safeguard system, and that safeguard system allows that same intelligence to be shipped and used by everyone with Fable 5. When a request first touches cybersecurity, biology, or chemistry topics, Fable will route it to our next most capable model, Opus 4.8. The response is clearly labeled and you're charged Opus prices. We know this is not perfect yet — researchers doing legitimate work in these fields will sometimes hit a block and reroute, and we're continuing to work on that. Our customers are already finding value. Rakuten found that at the highest effort levels Fable 5 reflects and validates its work, and for them that makes autonomous automated operations possible at scale. Cognition ran Fable 5 against their frontier coding eval and it scored the highest of any model they've tested. And Jensen Spark told us that Fable 5 came out number one in their evals, winning head-to-head against every model they've tested, strongest on the hardest tasks such as UI design and game coding.
So what about Mythos 5? It's the same underlying model as Fable 5, but with the cyber and bio safeguards lifted. It's available today for our Project Glasswing partners. Later this month, we'll begin enrolling more researchers from life sciences to access Mythos 5, because those same capabilities that make biology risky are the same ones that have the highest impact for real good with AI. As developers, what can you do with Fable 5? One metric I like to look at is time horizon, which is how long a model can work autonomously before losing coherence on what to do next. With Fable 5, we've seen agents that are proactive and know what to do without being told. Instead of asking Claude to write a project update, you can now ask Fable to make sure the project stays on track for the whole week. Instead of asking Claude to produce a financial forecast, you can tell Fable to own, update, iterate, and improve on the forecast to keep it accurate.
As developers, how do you take advantage of all this change? First, you should design for the next version of Claude, not just the current one. The developers who win are the ones whose architectures, harnesses, and product experiences are ready to absorb the next jump in intelligence. As models become more intelligent, they can dig farther with more basic primitives such as a file system or sandbox computing environments. You'll also need to start making harder eval prototypes for experiences that may not work yet. When a task or prototype that wasn't quite working starts working, that's a signal to ship something magical. And finally, the teams who win treat model upgrades as business opportunities — automated evals, testing processes, staying hands-on. We can't wait to see what you'll build next with Fable 5. And now, Angela and Caitlyn are going to show you a bit of how the Claude platform can make this a reality. Thank you so much.
Please welcome head of product for the Claude platform, Angela Jen. Last night, while we were all asleep, a product somewhere noticed that it was broken. This product read its own error reports, found the bug, wrote the fix, and rolled it out to all the users. By the time the team woke up, the problem was gone and the change log was already written. There was no standup, no ticket. This is what an AI-native company looks like. This is not a company where people just use AI to do their work, but a company where work itself runs on the substrate of AI and people decide what the outcome should be. This story can be a reality if you have the right ingredients, and there's three things to that. The first is the harness, the second is the context, and the third is the infrastructure. These three are what turn raw intelligence into true business outcomes, and that's exactly what the Claude platform provides. Recently, we packaged all these together through Cloud Managed Agents, our new product offering. Fable 5 is the best model for building long-running agents, and managed agents is purpose-built for Claude.
Let's talk about the harness. The harness is what gives Claude models the ability to actually do work. It includes tools, an environment, and the permission to act. You don't want AI that just gives you helpful suggestions; you want AI that can actually do the work. With Cloud Managed Agents, we offer an agentic harness that separates the brain from the hands. The brain decides what to do, and then we spin up sandboxes or hands to execute the work as needed. Our harness also has an iterative ability to operate on an outcome. You specify what that outcome is, and then a managed agent is able to iterate until it actually achieves it. Second, context. A model is only as good as the context you give it. We offer a 1 million context window so you can have agents that consume high amounts of content without any degradation. We also offer our agents memory so they can remember what they've been doing. We give our agents the ability to read and write skills of their own so they can fill in the knowledge gaps they're missing. And lastly, we give our agents the ability to dream so they can inspect their own previous trajectories and identify how to self-improve. And lastly, infrastructure. If you build really long-running autonomous agents, this requires extreme scale and reliability. Cloud Managed Agents automatically spins up and down sandboxes and has the ability to generate multiple agentic fleets when needed. Notion used managed agents to power agent orchestration directly within their product, so their users can delegate complex long-running work to Claude inside their workspace. Asana used managed agents to build AI teammates that work alongside humans inside Asana projects.
Now I'd love to welcome Caitlyn back. Back in February, Claude became the official thinking partner of the Atlassian Williams Formula 1 racing team. To show you an example of Cloud Managed Agents in action, we worked with a fictional racing team called Shankiro Racing, and we helped them build a dashboard to analyze their car. Here we have the dashboard. On the side we've got four research projects, each backed by a Cloud Managed Agent: aerodynamics, tire temperature, power unit, and driver safety. The goal for each project is for our agents to figure out what needs to change to better optimize our car. We've used the Cloud Managed Agents feature called outcomes, which means when an agent is doing its work, we can use a rubric to define what good looks like, and the agent can iterate until it does its work incredibly well. We can kick off any one of these agents — in this case, driver safety. What's cool about this dashboard being fully powered by AI is we can ask a question like how much does a car currently weigh and how is this weight distributed, and the dashboard will refresh with this information.
If I were a developer building this dashboard, I would use the Claude developer console as my home base for developer productivity, observability, and tooling. We have all our sessions from past runs, and we get really rich observability for everything the agent is doing — its responses, its tool calls, and when it finished. Starting today, you can even schedule these agents to run at any cadence you want. We can click create deployment, run a nightly safety check using our safety agent, and choose a schedule such as every morning at 9 a.m. But it's not enough to have agents that are easy to build and run; they also need to be powerful. First, we have memory: our agents can choose, as they're working, to write their findings to a file system, and look back into it in the future. On top of this, we've built a feature called dreaming. If memory is real-time deciding to write to a file system, dreaming allows the agent to look back over all of its past sessions and update its memory and its skills so it does even better next time. If we come back to Shankiro's dashboard, we can hit this button that says dream, and this kicks off a session for a new agent to look back on all the old sessions and write its findings to memory and skills. Everything you just saw — outcomes, scheduled deployments, and dreaming — is available today on the Cloud Platform.
Please welcome head of product for Claude Code, Cat Woo. Caitlyn and Angela just showed you how to build production agents on the Cloud Platform. With Claude Code, we're bringing that same leverage to your work as a developer — not agents you ship to customers, but agents that ship code for you. The mission of Claude Code is to bridge the difference between an idea and a shipped product. We build tools that elicit the frontier intelligence from our models and make them accessible to every builder. We don't think of ourselves as having a finished road map; we think of ourselves more like mountaineers climbing alongside you. I remember just last year I would give Claude Code a task and review every single edit. Now, many of us are using auto mode to delegate permissions to Claude and only checking in after Claude Code has tested its changes and put up a PR ready for review. Claude Code started in the CLI, and this is still the place for power users who want a minimal text interface and the most control. Then we added the IDE for users who want to follow along with all the code changes. Then we heard you're running multiple Claude Codes in parallel, which you've affectionately called multi-clauding, and we've added two new interfaces. The one I use most frequently is Claude Code on Cloud Desktop, designed for people who want a full-screen graphical interface with built-in previews, a sidebar control plane, and the ability to render images and rich outputs. Next is our newest surface, Claude agents view in the CLI, a control plane without having to leave the terminal. The VS Code IDE extension and the Claude Code on Cloud Desktop app are built on the Claude agent SDK, the same one many of you are building on today. At Anthropic, engineers on average ship 8x more code than they did in past years, even as the engineering team has grown.
Here's some feedback we've heard and the products we built. We heard you want to spend less time on code review, so we shipped a code review product that deploys a team of agents to catch critical bugs in your PRs. We heard you want to code on the go, so we launched remote control and Claude Code on iOS and Android. We heard you want to run Claude Code on new tickets, so we built routines: configure it once and it'll run on a schedule, a webhook, or an API call. We heard your security teams are having a hard time keeping up, so we built Claude security, which scans your codebase overnight and flags vulnerabilities, including the most critical ones. And finally, we heard you want better ways to run ambitious large-scale tasks like major refactors and migrations, so we launched dynamic workflows, which lets you kick off Claude Code to run in parallel across tens or hundreds of agents in a deterministic structure. Spotify uses Claude to migrate thousands of repos — they're merging over a thousand PRs a month into production and cut migration time down by 90%. At Mercari, the entire engineering team runs on Claude Code and they've measured that engineering output is up 90% year-over-year.
Let me show you one of my favorite new features in Claude Code. Let's imagine we're Kaizen Operations and we create apps for engineering teams. Right now our marketing website is in English only, but we want to launch to 13 additional markets. Let's start with how we'd normally do this with one Claude Code agent. We'll run this prompt asking Claude to convert our website to Japanese. Claude explores the codebase, creates a language picker, and translates all the text to Japanese. It took Claude about three minutes to complete this one translation, and we have 12 more to go. If we did this one at a time sequentially, it would take almost an hour. So instead we'll use a dynamic workflow, which lets Claude create a repeatable process and run each of the new translations in parallel. We'll prompt Claude to use a workflow for our translation to these 12 new languages. Claude creates a workflow, and we can open it in the side pane and watch as it works. You can see all 12 translation agents running at the same time. After they complete, it's going to create 12 more agents to verify its work across these new languages. This work would have previously required us to run 12 separate tasks, and it can now be done in just one prompt. We can also save this workflow as JavaScript code and reuse it. Claude has now created 12 new versions of the website, all localized with one action. We can use dynamic workflows for large-scale migrations, codebase audits, or performance optimizations — any large job that requires running many agents at once in a deterministic structure.
Everything you just saw is available today, including dynamic workflows and Claude Code in the Cloud Desktop app. And our newest model, Claude Fable 5, is available to all Claude Code users. This is really what every talk today was pointing at. Diane's capability curve, Angela's agents that run on infrastructure you control, and what I just showed you — these are three layers to one story. The remaining gap is just how fast we can put these great capabilities to work for us. I encourage you to spend the rest of today exploring these layers. Join research talks if you want to learn more about the latest model capabilities, join Cloud Platform sessions if you're building your own agents, or join Claude Code workshops if you want to bring Claude Code into your day-to-day development work. All of this runs on Claude Fable 5, the best model we've ever shipped for agentic work, and it's live today. Thank you all and enjoy Code with Claude.
TL;DR
- The Code with Claude Tokyo 2026 keynote announces Anthropic's fifth-generation models, Claude Fable 5 and Claude Mythos 5, their most capable models ever, and frames Anthropic as a platform company that helps developers close the gap between exponential model capability and linear business adoption.
- Speakers cover three layers — models (Diane), the Cloud Platform and Cloud Managed Agents (Angela and Caitlyn), and Claude Code (Cat) — with new managed-agent features (scheduled deployments, environment-variable vaults, memory, dreaming) and Claude Code surfaces (desktop app, agents view, dynamic workflows).
- Customer stories (Rakuten, Canva, Notion, Asana, Spotify, Mercari, Williams F1) illustrate AI-native companies where work runs on AI and people decide outcomes, and developers are urged to design for the next version of Claude, not just today's.
Takeaways
- Anthropic released its fifth-generation models, Claude Mythos 5 and Claude Fable 5; Fable 5 is the most capable generally available model and is state-of-the-art on nearly all tested AI benchmarks, while Mythos 5 is the same model with cyber and bio safeguards lifted, available to Project Glasswing partners.
- Anthropic positions itself as a platform company: year-over-year API volume is up nearly 17x, and the platform gives developers tools to build agents because most people will experience AI through products others build on it.
- Model capability is on an exponential while most business capability is linear, creating a growing gap that the Claude platform aims to close; the keynote traces progress from drafting commit messages to Mythos finding a 27-year-old OpenBSD vulnerability.
- Fable 5 excels at long-horizon autonomy (running for days on one goal across millions of tokens, dispatching sub agents), single-shot correctness, reading code, vision, and non-coding knowledge work; a new safeguard routes cyber/bio/chem requests to Opus 4.8.
- Cloud Managed Agents combine three ingredients — an agentic harness (separating the "brain" from sandbox "hands"), context (1M-token window, memory, self-written skills, "dreaming"), and production infrastructure (auto-scaling sandboxes, agentic fleets).
- New managed-agent features shipping today include scheduled deployments (run agents on any cadence) and environment-variable vaults for secure authenticated API requests, demonstrated on a fictional Shankiro racing dashboard with outcomes-based rubrics.
- Claude Code now spans the CLI, IDE, the Cloud Desktop app, and a new agents view, supporting "multi-clauding" in parallel; Anthropic engineers ship 8x more code than in past years, and products include code review, remote control on iOS/Android, routines, Claude security, and dynamic workflows.
- A dynamic-workflows demo localizes a marketing site into 12 languages in parallel (then verifies each), versus nearly an hour sequentially; customers like Spotify (1,000+ PRs/month, migration time cut 90%) and Mercari (engineering output up 90% YoY) show org-scale adoption.
Vocabulary
Claude Fable 5 — Anthropic's most capable generally available fifth-generation model, state-of-the-art on nearly all tested benchmarks.
Claude Mythos 5 — The same underlying model as Fable 5 with cyber and bio safeguards lifted, available to Project Glasswing partners.
Cloud Managed Agents — Anthropic's product for building and deploying long-running agents with a harness, context tools, and production infrastructure.
Harness — The tools, environment, and permissions that let a model actually do work, separating the deciding "brain" from the executing sandbox "hands."
Dreaming — A managed-agent feature where an agent reviews past sessions to update its memory and skills and self-improve.
Project Glasswing — Anthropic's program for safely releasing high-capability models, with a safeguard system routing sensitive requests.
Dynamic workflows — A Claude Code feature that runs Claude in parallel across many agents in a deterministic structure for large-scale tasks.
Time horizon — How long a model can work autonomously before losing coherence on what to do next.
AI-native company — A company where work itself runs on AI and people decide what the outcomes should be.
Transcript
Please welcome to the stage head of engineering for the Claude platform at Anthropic, Caitlyn Les. Tokyo, this is the first time that we've brought Code with Claude to Japan and we're grateful to spend the next couple of days with all of you. This morning we'll be talking about our models, our platform, and our products. But before we get into it, I want to start by sharing that just a few hours ago, we released the fifth generation of Claude models, Claude Mythos 5 and Claude Fable 5. These are our two most capable models ever. Diane, our head of product for research, will join us shortly to share a lot more about why these models are so special.
I lead engineering for the Claude platform. The platform gives developers the tools they need to build systems on top of Claude to harness its intelligence. This is the highest-leverage way for us to help solve the world's most important problems. And this is why Anthropic is a platform company. Developers all over the world, many in this room today, produce far more value on top of the platform than we could ever build on our own. So let's start with what I'm seeing from our customers lately. There's an incredible volume of powerful applications shipping right now, and a lot of it is coming from right here in Japan. Rakuten is one of our favorite customers to work with. Their team went from using Claude Code to accelerate development to building on Cloud Managed Agents to power custom internal agents across engineering, product, sales, and finance. One of their product managers coordinates teams of agents exactly the same way a leader manages teams of humans. Now they're shipping major releases every two weeks instead of only once per quarter. And it's the same across Asia-Pacific. Another great customer is Canva, the Australian design platform that hundreds of millions of people use. Most of their users have never written a line of code. But with the help of Claude, Canva Code changes all of this. Within a design, you can just ask for something interactive like a map, a calculator, or a widget, and Claude builds it into a working mini app ready to drop right into your page.
All around the world, people are building new systems and applications powered by Claude that nobody could build before. The landscape is changing faster than ever. Things that weren't possible yesterday have become possible today. We're on a mission to keep raising the ceiling of what's possible to build by making models that are increasingly capable. A couple of years ago, the frontier of model development was just the ability to draft a really simple commit message. One year ago, we were standing on stage at our first-ever Code with Claude event. Opus 4 was the headline, and it was mind-blowing at the time that Claude could build an entire feature on its own. Six months ago, agents were able to run overnight to complete long-running and autonomous tasks. Two months ago, Mythos read the entire OpenBSD source tree and found a 27-year-old vulnerability that had slipped past human reviews and static analysis for almost three decades. And earlier today, we released Mythos 5 and Fable 5. Fable 5 is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, scientific research, vision, and more. These jumps keep getting bigger, but the time intervals keep getting shorter. Even though these model capabilities are improving on this exponential, most business capabilities are still on a linear. So there's a growing gap between what AI can do and what it's actually doing for people. Closing that gap is our collective opportunity, and that is why we've built the Claude platform. Year-over-year API volume is up nearly 17 times on the platform.
This morning, Diane will talk about our models. Next, Angela and I will walk you through how you can build and deploy agents at scale on the Cloud Platform using Cloud Managed Agents. Today we're excited to share that we're shipping brand new Cloud Managed Agents features. You can now schedule deployments to have your agents run on whatever cadence you need, and you can store environment variables in vaults so your agents can securely make authenticated API requests without giving them access to the keys. Finally, Cat will walk you through the latest in Claude Code, where the average developer is now spending 20 hours per week with Claude. Most people will never integrate with the Claude API themselves. They'll experience AI through something that one of you built on the Claude platform, like a salesperson walking into a high-stakes meeting fully briefed by Slack agents, or a lawyer getting a brief out the door faster than ever with Legora. This is why we're a platform company. Thank you for being here, for partnering with us, and for showing us what's possible. Arigato gozaimasu.
Please welcome Diane from our research product team. Hi, good morning Tokyo. I'm Diane and I joined Anthropic in 2023, and I've been a part of every version of Claude since Claude 2. That's bringing 21 versions of Claude across Haiku, Sonnet, Opus, and now Fable and Mythos to end users and developers like you. Our most recent launch happened just a few hours ago. We released Claude Fable 5 and Claude Mythos 5, the first generation of our fifth models. Fable 5 is the most capable model we've ever made generally available. It's based on the same foundations as Mythos 5. We talk a lot about the exponential at Anthropic. As model intelligence increases, we believe the value of use cases it creates increases exponentially. For example, agentic coding that many of you are using today is far more valuable than simple autocomplete from just a few years ago.
Let's start with coding, because we believe that's where most of you will feel the magic first. Fable 5 performs the highest on SWE-bench Pro. But the benchmark itself undersells the story. The longer and more complicated and sophisticated the task, the farther the gap between Fable and every other model out there. Two things drive that gap. The first is single-shot correctness. Give Fable 5 a complex, well-specified problem and it will nail it on the first pass. Early testers have told us about single prompts that essentially create work that would have taken days or even weeks for a group of teams. The second is long-horizon autonomy. Fable 5 can run for days on a single goal and stay coherent the entire way. It'll remember your specifications even on tasks that span millions of tokens. And it can dispatch sub agents, keep them on track far more dependably and with more cost consciousness than any other model we've shipped. One more thing: Fable 5 isn't just good at writing code, it's even better at reading it. It's better at triaging outages and digging through your repo history to figure out what's broken when, and proactively surfacing suggestions for making it better. Fable 5 is just as much of a step outside of coding. Whether that's financial analysis, documents, slides, spreadsheets, Fable 5 manages that work end to end. It'll follow instructions, stay on scope, and what it creates back for you will be professional grade. It's also better where requests aren't clean. Fable 5's vision is also the best in the industry. It can read dense technical images, web applications, plots, diagrams, and charts far more accurately than any other version of Claude.
Now, intelligence of this magnitude cuts both ways. Two months ago, we launched Project Glasswing and made Mythos preview available to a small group of partners because its capabilities in cybersecurity were strong enough to be potentially misused. Since then, we built a new safeguard system, and that safeguard system allows that same intelligence to be shipped and used by everyone with Fable 5. When a request first touches cybersecurity, biology, or chemistry topics, Fable will route it to our next most capable model, Opus 4.8. The response is clearly labeled and you're charged Opus prices. We know this is not perfect yet — researchers doing legitimate work in these fields will sometimes hit a block and reroute, and we're continuing to work on that. Our customers are already finding value. Rakuten found that at the highest effort levels Fable 5 reflects and validates its work, and for them that makes autonomous automated operations possible at scale. Cognition ran Fable 5 against their frontier coding eval and it scored the highest of any model they've tested. And Jensen Spark told us that Fable 5 came out number one in their evals, winning head-to-head against every model they've tested, strongest on the hardest tasks such as UI design and game coding.
So what about Mythos 5? It's the same underlying model as Fable 5, but with the cyber and bio safeguards lifted. It's available today for our Project Glasswing partners. Later this month, we'll begin enrolling more researchers from life sciences to access Mythos 5, because those same capabilities that make biology risky are the same ones that have the highest impact for real good with AI. As developers, what can you do with Fable 5? One metric I like to look at is time horizon, which is how long a model can work autonomously before losing coherence on what to do next. With Fable 5, we've seen agents that are proactive and know what to do without being told. Instead of asking Claude to write a project update, you can now ask Fable to make sure the project stays on track for the whole week. Instead of asking Claude to produce a financial forecast, you can tell Fable to own, update, iterate, and improve on the forecast to keep it accurate.
As developers, how do you take advantage of all this change? First, you should design for the next version of Claude, not just the current one. The developers who win are the ones whose architectures, harnesses, and product experiences are ready to absorb the next jump in intelligence. As models become more intelligent, they can dig farther with more basic primitives such as a file system or sandbox computing environments. You'll also need to start making harder eval prototypes for experiences that may not work yet. When a task or prototype that wasn't quite working starts working, that's a signal to ship something magical. And finally, the teams who win treat model upgrades as business opportunities — automated evals, testing processes, staying hands-on. We can't wait to see what you'll build next with Fable 5. And now, Angela and Caitlyn are going to show you a bit of how the Claude platform can make this a reality. Thank you so much.
Please welcome head of product for the Claude platform, Angela Jen. Last night, while we were all asleep, a product somewhere noticed that it was broken. This product read its own error reports, found the bug, wrote the fix, and rolled it out to all the users. By the time the team woke up, the problem was gone and the change log was already written. There was no standup, no ticket. This is what an AI-native company looks like. This is not a company where people just use AI to do their work, but a company where work itself runs on the substrate of AI and people decide what the outcome should be. This story can be a reality if you have the right ingredients, and there's three things to that. The first is the harness, the second is the context, and the third is the infrastructure. These three are what turn raw intelligence into true business outcomes, and that's exactly what the Claude platform provides. Recently, we packaged all these together through Cloud Managed Agents, our new product offering. Fable 5 is the best model for building long-running agents, and managed agents is purpose-built for Claude.
Let's talk about the harness. The harness is what gives Claude models the ability to actually do work. It includes tools, an environment, and the permission to act. You don't want AI that just gives you helpful suggestions; you want AI that can actually do the work. With Cloud Managed Agents, we offer an agentic harness that separates the brain from the hands. The brain decides what to do, and then we spin up sandboxes or hands to execute the work as needed. Our harness also has an iterative ability to operate on an outcome. You specify what that outcome is, and then a managed agent is able to iterate until it actually achieves it. Second, context. A model is only as good as the context you give it. We offer a 1 million context window so you can have agents that consume high amounts of content without any degradation. We also offer our agents memory so they can remember what they've been doing. We give our agents the ability to read and write skills of their own so they can fill in the knowledge gaps they're missing. And lastly, we give our agents the ability to dream so they can inspect their own previous trajectories and identify how to self-improve. And lastly, infrastructure. If you build really long-running autonomous agents, this requires extreme scale and reliability. Cloud Managed Agents automatically spins up and down sandboxes and has the ability to generate multiple agentic fleets when needed. Notion used managed agents to power agent orchestration directly within their product, so their users can delegate complex long-running work to Claude inside their workspace. Asana used managed agents to build AI teammates that work alongside humans inside Asana projects.
Now I'd love to welcome Caitlyn back. Back in February, Claude became the official thinking partner of the Atlassian Williams Formula 1 racing team. To show you an example of Cloud Managed Agents in action, we worked with a fictional racing team called Shankiro Racing, and we helped them build a dashboard to analyze their car. Here we have the dashboard. On the side we've got four research projects, each backed by a Cloud Managed Agent: aerodynamics, tire temperature, power unit, and driver safety. The goal for each project is for our agents to figure out what needs to change to better optimize our car. We've used the Cloud Managed Agents feature called outcomes, which means when an agent is doing its work, we can use a rubric to define what good looks like, and the agent can iterate until it does its work incredibly well. We can kick off any one of these agents — in this case, driver safety. What's cool about this dashboard being fully powered by AI is we can ask a question like how much does a car currently weigh and how is this weight distributed, and the dashboard will refresh with this information.
If I were a developer building this dashboard, I would use the Claude developer console as my home base for developer productivity, observability, and tooling. We have all our sessions from past runs, and we get really rich observability for everything the agent is doing — its responses, its tool calls, and when it finished. Starting today, you can even schedule these agents to run at any cadence you want. We can click create deployment, run a nightly safety check using our safety agent, and choose a schedule such as every morning at 9 a.m. But it's not enough to have agents that are easy to build and run; they also need to be powerful. First, we have memory: our agents can choose, as they're working, to write their findings to a file system, and look back into it in the future. On top of this, we've built a feature called dreaming. If memory is real-time deciding to write to a file system, dreaming allows the agent to look back over all of its past sessions and update its memory and its skills so it does even better next time. If we come back to Shankiro's dashboard, we can hit this button that says dream, and this kicks off a session for a new agent to look back on all the old sessions and write its findings to memory and skills. Everything you just saw — outcomes, scheduled deployments, and dreaming — is available today on the Cloud Platform.
Please welcome head of product for Claude Code, Cat Woo. Caitlyn and Angela just showed you how to build production agents on the Cloud Platform. With Claude Code, we're bringing that same leverage to your work as a developer — not agents you ship to customers, but agents that ship code for you. The mission of Claude Code is to bridge the difference between an idea and a shipped product. We build tools that elicit the frontier intelligence from our models and make them accessible to every builder. We don't think of ourselves as having a finished road map; we think of ourselves more like mountaineers climbing alongside you. I remember just last year I would give Claude Code a task and review every single edit. Now, many of us are using auto mode to delegate permissions to Claude and only checking in after Claude Code has tested its changes and put up a PR ready for review. Claude Code started in the CLI, and this is still the place for power users who want a minimal text interface and the most control. Then we added the IDE for users who want to follow along with all the code changes. Then we heard you're running multiple Claude Codes in parallel, which you've affectionately called multi-clauding, and we've added two new interfaces. The one I use most frequently is Claude Code on Cloud Desktop, designed for people who want a full-screen graphical interface with built-in previews, a sidebar control plane, and the ability to render images and rich outputs. Next is our newest surface, Claude agents view in the CLI, a control plane without having to leave the terminal. The VS Code IDE extension and the Claude Code on Cloud Desktop app are built on the Claude agent SDK, the same one many of you are building on today. At Anthropic, engineers on average ship 8x more code than they did in past years, even as the engineering team has grown.
Here's some feedback we've heard and the products we built. We heard you want to spend less time on code review, so we shipped a code review product that deploys a team of agents to catch critical bugs in your PRs. We heard you want to code on the go, so we launched remote control and Claude Code on iOS and Android. We heard you want to run Claude Code on new tickets, so we built routines: configure it once and it'll run on a schedule, a webhook, or an API call. We heard your security teams are having a hard time keeping up, so we built Claude security, which scans your codebase overnight and flags vulnerabilities, including the most critical ones. And finally, we heard you want better ways to run ambitious large-scale tasks like major refactors and migrations, so we launched dynamic workflows, which lets you kick off Claude Code to run in parallel across tens or hundreds of agents in a deterministic structure. Spotify uses Claude to migrate thousands of repos — they're merging over a thousand PRs a month into production and cut migration time down by 90%. At Mercari, the entire engineering team runs on Claude Code and they've measured that engineering output is up 90% year-over-year.
Let me show you one of my favorite new features in Claude Code. Let's imagine we're Kaizen Operations and we create apps for engineering teams. Right now our marketing website is in English only, but we want to launch to 13 additional markets. Let's start with how we'd normally do this with one Claude Code agent. We'll run this prompt asking Claude to convert our website to Japanese. Claude explores the codebase, creates a language picker, and translates all the text to Japanese. It took Claude about three minutes to complete this one translation, and we have 12 more to go. If we did this one at a time sequentially, it would take almost an hour. So instead we'll use a dynamic workflow, which lets Claude create a repeatable process and run each of the new translations in parallel. We'll prompt Claude to use a workflow for our translation to these 12 new languages. Claude creates a workflow, and we can open it in the side pane and watch as it works. You can see all 12 translation agents running at the same time. After they complete, it's going to create 12 more agents to verify its work across these new languages. This work would have previously required us to run 12 separate tasks, and it can now be done in just one prompt. We can also save this workflow as JavaScript code and reuse it. Claude has now created 12 new versions of the website, all localized with one action. We can use dynamic workflows for large-scale migrations, codebase audits, or performance optimizations — any large job that requires running many agents at once in a deterministic structure.
Everything you just saw is available today, including dynamic workflows and Claude Code in the Cloud Desktop app. And our newest model, Claude Fable 5, is available to all Claude Code users. This is really what every talk today was pointing at. Diane's capability curve, Angela's agents that run on infrastructure you control, and what I just showed you — these are three layers to one story. The remaining gap is just how fast we can put these great capabilities to work for us. I encourage you to spend the rest of today exploring these layers. Join research talks if you want to learn more about the latest model capabilities, join Cloud Platform sessions if you're building your own agents, or join Claude Code workshops if you want to bring Claude Code into your day-to-day development work. All of this runs on Claude Fable 5, the best model we've ever shipped for agentic work, and it's live today. Thank you all and enjoy Code with Claude.