AWS Open-Sources Strands Decider 2B, a 2B Agent Model Scoring 72.3% on JevBench With No Text Output

News Summary
Amazon Web Services announced Strands Decider 2B on Thursday, October 1, 2026 (Eastern Time), an open-source decision model that cannot write a single word of text. Released through the Strands Labs program, the roughly 2-billion-parameter model is built to make the small, routine choices that AI agents face constantly, such as picking a tool, routing a request or classifying a policy question, while larger language models are saved for harder reasoning. The weights, training data sources and scripts are public under the Apache 2.0 license.
What Strands Decider 2B Is
Strands Decider 2B is a specialized "decision model" rather than a chatbot. Instead of generating text, it picks from a list of options supplied by the developer, assigns a score, or returns a confidence estimate. AWS describes it as a small, fast model optimized for quick experimentation and local development in agentic AI. The release was added to strands-labs, the experimental arm of the Strands Agents ecosystem. Coverage describes it as a "system one" component: fast, intuitive judgment that sits inside an agent's control loop.
The Strands Agents SDK itself was released roughly 16 months before this announcement, and the project has since grown to cover much of what an agent needs, including a harness released in September 2026. Decider 2B adds a small, local piece to that stack.
How It Works
According to reporting by Tech Times and The Letter Two, AWS started from Alibaba's open Qwen3.5-2B base model (about 1.9 billion parameters). The team fine-tuned it with LoRA, a lightweight technique that trains a small add-on layer rather than rewriting the whole network. They then removed the language-modeling head, which is what normally lets a model produce text, and replaced it with a "pointer head" of about one million parameters.
The pointer head compares the model's internal representation of the input against each candidate option using a dot product, then applies a masked softmax to produce a probability distribution. Because of this design, the output is constrained entirely to the options provided. The model supports three question types: yes/no, choice among N candidates, and score placement on an ordered rubric. LoRA was applied at rank 16, and the pointer head runs in fp32 precision.
Benchmarks and Latency
On the public JevBench v1 evaluation, Decider 2B answered 167 of 231 tasks correctly, an accuracy of 0.723. Reported calibration figures include a Brier score of 0.342 and an expected calibration error of 0.052. Performance varied by difficulty: near-perfect on easy tasks, about 87.5% on standard tasks and about 50.5% on hard tasks.
Rankings differ slightly between outlets. Tech Times reports a third-place finish among 33 models in the 2B class, and first among 30 models once larger ones are excluded. The Letter Two describes it as second among public models of similar size and first among those that publish a full training recipe. Readers should treat the exact rank as dependent on how the leaderboard is filtered.
On speed, AWS says decisions run in tens of milliseconds, under 100 milliseconds on widely available hardware. Independent figures in Tech Times are more conservative: a median of 115 ms (95th percentile 299 ms) on an Nvidia RTX 3090, and a warm median of 153 ms on an Apple M3 Pro for inputs under 300 tokens. Asking two questions about the same input costs roughly 230 ms. These community numbers were not verified by AWS.
Training and Reproducibility
A distinctive point of the release is openness. AWS published the training data sources and licenses, the training scripts and evaluation tooling. Tech Times reports that training takes about 11 hours on a single RTX 3090 with 24 GB of memory, or about 1 hour 10 minutes on eight H100 GPUs. The project also documents predictions and failure conditions before each training run, a pre-registration style practice that is uncommon in model releases.
Availability and Use Cases
The model is available on Hugging Face (StrandsAgents/strands-decider-2B-hobson-v19) and the code lives on GitHub under strands-labs/strands-decider. It installs with pip install strands-decider and includes a command-line tool and a local HTTP server. The server listens on localhost with no authentication by default, so teams should add their own access controls before exposing it. It can run on a CPU, a consumer GPU or Apple silicon, and works offline without an API dependency.
Suggested uses include model routing, tool selection, memory and context management, policy classification, guardrails, evaluations and output validation. In a hybrid design, a large model handles complex reasoning while Decider 2B handles the high-volume routine decisions cheaply and locally.
Competitive Context
The release arrives amid a small wave of decision-model projects. TypeSafe AI's Jev model, along with LocalJev and OpenJev, appears on the public jevbench.dev leaderboard, and Cloudflare released its Clef and Clef-flash models the same week. Some commentators framed the AWS move as undercutting commercial offerings, while AWS itself positions the project as an experiment in whether specialized models can handle routine agent work.
Why It Matters for Learners
For students and developers, the project is a clear example of a broader engineering idea: not every step in an AI system needs a large generative model. Constraining a model's output to a fixed set of options makes it faster, easier to test and easier to reason about. Because the code, data sources and scripts are open, it is also a practical, reproducible case study in LoRA fine-tuning and calibration measurement.
Caveats
Hard-task accuracy of about 50% shows the model is not a replacement for deeper reasoning. Latency claims vary by hardware and by source, and some benchmark figures come from community testing rather than AWS. Exact timestamps of the announcement were not given in the sources reviewed; dates above follow Eastern Time as reported.