CoreWeave Launches Forge to Connect AI Deployment, Monitoring and Improvement
The platform brings production feedback into training and evaluation. Pro starts at $60 a month, but teams with 50 or more employees must use Enterprise.
Listen to this story
The audio brief
Story brief
3 key pointsForge packages CoreWeave’s former Weights & Biases, OpenPipe and Marimo products with new services in a shared environment for improving deployed AI systems. Its core workflow carries production traces into datasets, training and release evaluations, so teams can investigate failures and test fixes without relying on disconnected handoffs. The offer is integrated with CoreWeave’s cloud but does not require a cloud move; pricing and eligibility are tiered, with Pro limited to teams under 50 employees.
- 01
Agent Lens records agent steps, decisions and tool calls; flagged failures become versioned datasets that retain production context.
- 02
Teams can improve prompts, tools or retrieval without retraining; training options include supervised fine-tuning, reinforcement learning and distillation.
- 03
Sandboxes can run agents and evaluations on infrastructure teams already use, while Forge also supports workloads on other clouds.
CoreWeave launched Forge on September 30, 2026, bringing AI deployment, monitoring, data preparation, training and evaluation into one development environment. Available starting today, the platform is designed to turn problems found in live models and agents into tested improvements. CoreWeave says teams can connect workloads on other clouds or their own infrastructure rather than move everything onto its cloud.
Announced at Fully Connected 2026, Forge addresses a problem CoreWeave describes as a disconnected improvement cycle. A production team may see whether a service is running, while researchers lack the conversations needed to judge answer quality. Evaluation tests can fall behind changing user behavior. The company’s pitch is to connect those records and workflows, reducing the manual handoffs between finding an issue and checking a fix.
A failed run becomes material for the next version
The connection starts with traces: records of an agent’s steps, decisions and tool calls. Agent Lens captures those records, and monitors score live behavior against baselines set by the team. Flagged failures become versioned datasets in Weights & Biases Models, retaining the production context. Those examples can then inform both training and the tests used to decide whether a replacement is ready to ship.
An improvement need not mean retraining a model. Forge’s workflow can begin with changes to a prompt, a tool or retrieval—the process of finding information for a model to use. When training is needed, CoreWeave offers supervised fine-tuning on examples, reinforcement learning using reward signals, and distillation. The company says teams can use these services without setting up a training cluster.
New services alongside familiar tools
Forge is not a collection of entirely new products. It brings together the former Weights & Biases, OpenPipe and Marimo products, alongside CoreWeave services. The launch introduces Agent Lens, Model Distillation and Notebooks as new services; ARIA and Sandboxes are now generally available. Three components illustrate the different jobs within that shared environment:
- Model Distillation trains a smaller, open-weights model on a larger model’s outputs, then compares it against the model currently serving the task.
- Notebooks provides managed Python workspaces for experiments, collaboration, evaluations and custom analysis of models and agents.
- Sandboxes runs agents, tool calls, reinforcement learning and evaluations in isolated CPU or GPU environments, including on infrastructure teams already use for training.
ARIA adds assistance across that work: it analyzes experiment history, proposes experiments and recommends code changes. CoreWeave describes it as advisory, with the user’s judgment remaining in charge. Evaluations compare a candidate with the current version using production traces. Registry records the dataset and model checkpoint that passed, plus the version to roll back to. The workflow therefore includes a check before deployment, not just suggestions for changes.
One account, without a required cloud move
CoreWeave says Forge works with teams’ existing models, frameworks and clouds, while keeping generated improvement data in portable formats. That openness sits alongside tighter integration under one account and navigation. CoreWeave also says the platform runs best on its own cloud. Teams can connect external workloads and choose CoreWeave’s training or inference services when they want the infrastructure as well.
The Forge pricing page lists Free at $0 a month and Pro starting at $60 a month, billed monthly, with a 30-day trial. Pro is restricted to early-stage teams with fewer than 50 employees; customers exceeding that limit must transition to custom-priced Enterprise plans. Existing Weights & Biases users keep their credentials, projects and wandb.ai links, with the other Forge tools accessible through the same account.
Story updates
Latest developments
CoreWeave Opens Agent Lens Preview to Find AI-Agent Failures and Check Fixes
Teams running AI agents can now try a tool designed to turn live failures into tests for the next version, rather than investigate each conversation separately. CoreWeave opened Agent Lens in public preview on September 30, 2026. Its launch announcement describes a connected workflow for finding recurring problems, diagnosing them and checking proposed fixes before release.
Agent Lens works from traces: records of an agent’s inputs, outputs, model responses and tool calls. Its conversation view presents those records as readable dialogue. Product managers and subject-matter experts can review what happened, while engineers can expand individual turns to inspect the execution details behind them.
CoreWeave says the system analyzes the full production collection rather than a small, manually selected sample. It groups conversations with the same underlying failure, even when their wording, tool sequence or visible error differs. Each group links back to source traces and representative examples, giving teams evidence to inspect rather than just a changed dashboard metric.
Those groups update as new traces arrive. Teams can rank problems by frequency, changes over time and affected workflows. Agent Lens also groups requests by user intent, showing which tasks customers want and where the agent falls short. CoreWeave presents the two views as complementary: a popular task may deserve investment, while a rarer failure may demand attention because of safety or compliance risk.
Finding a suspicious pattern is only the start. Agent Lens lets teams create custom LLM judges—language models used to score agent behavior—based on failures found in production. Domain experts define the expected response and score representative examples, establishing the standards the automated judge should apply.
The tool compares human labels with judge scores and surfaces disagreements. Teams can examine cases the judge wrongly flags or misses, clarify scoring rules and refine its criteria before applying it broadly. Once aligned with those standards, the judge can score new production traces automatically. The workflow therefore includes checking the evaluator, not simply accepting its first verdict.
Sources
- coreweave.comIntroducing CoreWeave Forge | CoreWeave Blog
- coreweave.comCoreWeave Forge | Your AI Loop, Connected
Reader comments
Newest comments first. Replies stay oldest first.