LangSmith adds agent red teaming and user-scoped memory: what production agents need next
Building an agent is the easy week. Keeping it honest in production is the rest of the year.
What happened
LangChain shipped two updates aimed at agents already in production.
LangSmith Engine v2 is an in-platform agent that reviews your agent's production traces. LangChain says Engine has analyzed more than 70 million traces since May. The new version adds Red Teaming, which studies your agent's traces and repo to test for weaknesses like hallucinations and system prompt violations before users hit them. It also flags inefficient agent paths, such as repetitive tool calls, and tracks trends in error rate, latency and cost. For agents on LangSmith Deployment, Engine now reproduces a failure, proposes a fix, tests it against the same inputs and only then puts it in your review queue. Red Teaming and auto-validated fixes are in private beta.
Managed Deep Agents 0.8 adds user-owned credentials, user-level memory, HTTP channels, file transfer in Slack and built-in web search. Memory now has two layers: agent memory shared across the deployment, and user memory keyed to the authenticated person, so one user's preferences do not leak into a group conversation. Credentials can be scoped at the user or agent level.
My take
Two features here solve problems I see in almost every client agent:
- Identity. An internal sales or support agent used by ten people should act with each person's permissions, not one shared admin key. User-owned credentials make that the default instead of a custom build.
- Silent cost creep. Agents that loop on the same tool call rarely throw errors. They just get slower and pricier. Tracking trajectory efficiency and cost trends catches that before the monthly bill does.
You do not need LangSmith to apply this. Whatever you run, n8n, Make or custom code, log every tool call, scope credentials per user, keep shared rules separate from personal memory, and review the slowest and most expensive runs every week. The testing-before-review step is the part I would copy first: a fix that has not been replayed against the input that broke it is still a guess.
More posts
- Meta gives every Muse user a cloud computer with credentials kept outside the agent. A blueprint worth copyingSep 27, 2026
- Google is testing a Buy button for Flipkart inside Gemini and AI Mode. Checkout is moving into the AI answerSep 27, 2026
- Having AI advice available nearly wiped out people's willingness to say I don't know, even when the AI was wrongSep 27, 2026
