When we published our first AI Technology Radar in July, we made a promise: it would age in public. If a recommendation changed, you would see the change and the reasoning behind it, not a quiet internal memo. This second edition is the first real test of that promise, and it is not an entirely comfortable one. Two of the most widely deployed tools on the radar moved backwards. Thirty entries left it altogether. The radar is smaller than it was in May.
That is the point. This post walks through what moved, what the evidence was, and why we changed the radar's own rules so that it is allowed to shrink.
The shape of this edition
The September 2026 edition came in two releases. The first, on 4 September, touched 42 of 153 entries. A second pass on 6 September, run against the result of the first, touched 27 more. Together they leave the radar at 136 items:
- 30 retired. Entries the radar no longer needs an opinion on.
- 7 moved rings. Four promotions to Adopt, one to Trial, and two demotions.
- 19 refreshed. Rewritten on new evidence without changing ring.
- 13 added. Twelve in Assess and one in Hold.
After both releases the rings hold 26 Adopt (19%), 38 Trial (28%), 58 Assess (43%) and 14 Hold (10%). Assess is still the largest ring, and it should be. This is a field that keeps producing more candidates than proven practice. What is new is that the radar now has a way to stop carrying candidates that never turned into anything.
Two demotions on security evidence
The most visible change is that LangChain and MLflow 3, both in Adopt since our first edition, have moved to Trial. Neither is a recommendation to stop using them. Both are a change in what using them has to involve.
LangChain. In March, Cyera's coordinated "LangDrained" disclosure
documented three high- and critical-severity vulnerabilities across LangChain
and LangGraph: a serialization injection rated CVSS 9.3, a path traversal,
and a SQL injection in the SQLite checkpointer. Each one exposed a different
class of enterprise data, from filesystem contents to API credentials to
stored conversation histories[1].
The patches shipped. Then in May another high-severity advisory covered
unsafe deserialization through overly broad load() allowlists, affecting
both the 0.3 and 1.x lines[2].
The same class of defect came back. Trial means you keep building on it, but
only inside explicit guardrails around versions, document loaders,
serialization and checkpoint stores, and only if your team has the capacity
to track advisories and ship patch upgrades quickly.
MLflow 3. CVE-2026-64849 is a critical, unauthenticated server-side request forgery in MLflow's webhook test endpoint. It affects every version before 3.15.0. Because the endpoint reflects the upstream response back to the caller, a DNS-rebinding attacker can read cloud instance-metadata credentials straight out of the server. Attackers were exploiting it against exposed tracking servers within hours of the CVE being assigned, and CISA added it to the Known Exploited Vulnerabilities catalog[3]. And it is not a single-CVE story. Ten advisories between April and August described missing or misapplied authorization checks on MLflow endpoints. Trial here means patching to 3.15.0 or later, isolating the network, and never exposing a tracking server, treated as a hard gate rather than an item on the hardening backlog.
Both entries still describe tools with enormous production footprints and an intact value proposition. What earned them Adopt in May was maturity. What moved them in September is that running them safely is now a real operational commitment, and Adopt should not carry that asterisk quietly.
What earned Adopt
Four entries moved up to Adopt across the two releases, all on repeated production evidence. One of them is a control rather than a capability.
The NIST AI Risk Management Framework moved from Trial to Adopt. In May it was a well-regarded document. Since then NIST has published use cases and profiles, and organizations such as Workday have documented how they aligned their control frameworks to it[4]. That is the difference between a framework people cite and one people actually operate.
LLMOps platforms moved to Adopt as a category, backed by production case studies from ZenML and named LangSmith deployments at Schneider Electric, Rippling and Toyota. Dagster moved to Adopt on documented production use at US Foods, easyJet Holidays and PostHog[5].
The second release added the fourth. CrewAI moved from Trial to Adopt on named enterprise deployments at PwC, IBM and Gelato, and on AWS, where it underpins Bedrock agents[6]. The framework risks described further down are real, and the entry says so. They belong in production controls, not in gatekeeping.
AI platform engineering entered Trial on the strength of pilots from Nirmata, Pulumi Neo and Itential FlowAI.
Thirteen additions, twelve of them in Assess
The first release added six entries, all in Assess, and five of the six are about governing agents rather than building them:
- Microsoft's Agent Governance Toolkit, open-source runtime security controls for autonomous agents[7]
- asago, Red Hat's open-source project for turning governance policy into deployed, enforceable controls[8]
- The OWASP MCP Top 10, now a standalone entry because the Model Context Protocol has grown a risk surface of its own[9]
- Agent evaluation harnesses and agent trajectory evaluation, a body of work that is converging on measuring agent reliability as an engineering discipline rather than a benchmark score
The sixth, Rerun, is the outlier. It is a data layer for multimodal, multi-rate time series in robotics and physical AI, which is a gap our text-centric MLOps entries had not covered.
The second release added seven more. Six start in Assess: agent memory layers, agentic data control planes, agentic test automation, agentic vulnerability research for code, AI control protocol evaluation and Apache Fluss, the streaming storage layer for lakehouses that graduated to an Apache top-level project this year with production cases at Rednote and Taobao[10].
Agent memory layers deserves a word, because on 4 September we retired the standalone Mem0 entry for lack of signal. Two days later the second pass found the category rather than the product. Mem0, Zep, Cognee, Letta, TencentDB, Redis and Oracle now all ship persistent memory for agents, and Mem0's own report counts 21 frameworks and 20 vector stores integrated[11]. That is a distinct concern from retrieval, and it now has an entry of its own.
The seventh is the only addition that starts in Hold: agent framework supply chain risk. Check Point found nearly a dozen flaws across major agent frameworks this summer and argued that prompt injection is not the bug, the frameworks are[12]. Langflow produced the first agent-framework vulnerability on CISA's Known Exploited Vulnerabilities list, and a fully patched LiteLLM gateway was still hijackable by anyone holding an admin key[13]. The hold is not against agent frameworks. It is against adopting them without sandboxing, dependency control, least privilege and runtime monitoring.
The refreshed entries follow the same thread. Prompt injection defenses is now framed around permission bounding and tool-call monitoring. AI-augmented CI/CD leads with a documented attack class: prompt injection against agents that triage issues and review pull requests while holding elevated repository privileges[14]. And the AutoGen entry records that AutoGen is entering maintenance while Microsoft Agent Framework has reached 1.0 GA as its successor[15]. The second release rewrote ten more entries the same way. GitHub Copilot stays in Adopt but now carries its documented prompt-injection and exfiltration advisories, pgvector stays in Trial with an index-build CVE, and Milvus stays in Assess despite critical authentication-bypass advisories. The ring did not move. The threat model did.
Why we retired 30 entries
Here is the uncomfortable thing we learned while building this edition: the radar we designed in May could only grow. Every action available to us, whether adding, moving or refreshing, either kept an entry on the radar or put a new one there. There was no way to say that we no longer needed an opinion on something. Give that a few quarters and you get a radar where everything drifts towards Adopt, the Hold ring fills up with warnings nobody needs any more, and the list stops being a decision aid and turns into an inventory.
So we changed two rules.
Retirement is now an action. An entry can be taken off the radar without losing its history. Its write-up stays in the release it was published in, and a later edition can bring it back. Retiring something is not a judgement that it is bad. Hold says that, and a Hold entry earns its place by warning people off. Retiring simply means the radar no longer needs to carry the opinion.
Every quadrant has a cap of 30. A radar cannot have everything on it. Once a quadrant is full, adding an entry means removing one, and the review has to argue for which. Two of our four quadrants now sit exactly at the cap, AI & Data Engineering is at 31, and Developer AI & Delivery is at 45, down from 52. It will get there over the next two editions rather than in one cut, because we limit retirements to eight per quadrant per release so that each one can actually be reviewed.
Most of the 30 retirements are absorptions rather than rejections. A standalone agent-to-agent protocol entry, typed tool interfaces, KServe and several IDE agents were retired because the decision they informed is now covered by a broader entry. Serving now sits with vLLM, NVIDIA Dynamo and Kubernetes for AI inference. Typed tools fold into structured outputs. A few were Hold warnings that had simply run their course: standalone AutoML platforms, bespoke ETL scripts, homomorphic encryption in production AI, and blind LLM-as-judge scoring. The industry has moved on, and the warning no longer needs the space.
The second release retired nine more on the same rule: Bloom, federated learning, Figma Make, Hadoop HDFS, Monte Carlo, PageIndex, the Pi coding agent, Prefect 3 and Seldon Core v1. Each had gone unrevised since May and was absent from a harvest of more than 2,400 signals.
How we decided, and what we would not let the tooling decide
We use an agent to maintain this radar, and this edition is where we found its limits. We ran it five times against the same May baseline, asking the same question each time: what changed this quarter? It gave us five substantially different answers. Across all five runs it produced 162 distinct proposals, and only 24 of them showed up in more than one run.
We did not publish any of the proposals that appeared only once. Everything that moved or was added in this edition was proposed independently by at least two runs, or was a ring move from the run with the deepest evidence. The retirements were filtered the other way round. Anything one run wanted to retire while another wanted to promote, such as Feast, AI inference gateways and AI-augmented CI/CD, stayed on the radar. Absence of signal in one harvest is not evidence. Absence across several is.
The second pass was the same discipline applied to the tool's own gaps. Between the two releases we found that three of its nine search providers had been silently dead for months, that its search requests were not being counted against its budget, and that it read less than half of what it found. We fixed those, ran it three more times against the 4 September result, and again published only what two runs agreed on: 29 of 125 distinct proposals, 27 after our own review. We also started keeping a short list of technologies we know it should find. Its scouts reach three quarters of that list. Its slate holds a third. That gap is the next thing we will work on, and the December edition will say whether it closed.
That is the discipline the format demands. The tool finds the candidates. The evidence decides. And when the evidence says two of your Adopt entries need an asterisk, you print the asterisk.
You can explore every entry, ring and quadrant in the full RUBINLAKE AI Technology Radar, including the complete methodology. The next edition is due in December.

