What is Action-graded severity scale?
The Action-Graded Severity Scale ranks an agent's harmful actions from reversible mistakes to actions that cross boundaries or gain more power.
Field notes · 131 sourced records
AI agent field notes track the launches, tests, methods, and research ideas that may change how agents are built or judged. Each note links to the original source and shows when we checked it. Early findings stay provisional until better proof arrives.
The Action-Graded Severity Scale ranks an agent's harmful actions from reversible mistakes to actions that cross boundaries or gain more power.
Aggregating these outcomes yields an adversarial psychometric rating system that can scale with the systems being measured.
AgentBench is a benchmark that tests language-model agents across eight different interactive environments.
AgentLens is a benchmark that grades interactive coding agents across their whole working process, not only whether the final task passes.
ARC-AGI-2 is a benchmark of unfamiliar reasoning puzzles designed to test how well AI systems learn new rules.
Arize Phoenix is an open-source observability tool for tracing and evaluating AI applications, retrieval systems, and agents.
Audio Judge Reliability tests whether Gemini models can grade two-way voice-agent conversations closely enough to support, or sometimes replace, human reviewers.
An owner-published AI model leaderboard with methodology, evidence-status lanes, and an agentic evaluation category.
BioSecBench-Refusal is a benchmark for risk identification and refusal behavior for biological research tasks.
Braintrust helps teams test AI features, compare experiments, and trace what happened during a run.
CausalDS is a benchmark for evaluating causal reasoning in agentic data-science workflows.
ClassicLogic is a benchmark that tests whether an agent can learn problem-solving methods and combine them on unfamiliar tasks.
DataGovBench is a benchmark derived from governmental open data designed to evaluate LLMs in practical scenarios.
The Deliberative Collaboration Agent Benchmark tests whether several agents can share different pieces of information and reason together toward one decision.
Deployment Simulation for LLM Safety replays real conversation openings with a new model to estimate harmful behavior before that model reaches users.
Developer-Aligned Software Agent Evaluation is a method for judging coding agents on real development work, their full action history, contamination risk, and failures that matter to people.
Entroly selects and removes duplicate code context before an agent sees it, so the agent can work with a smaller prompt.
EvalLoop is a method for improving language-model applications through repeated evaluation, diagnosis, and repair.
EvoAgentBench is a benchmark for agent self-evolution via Ability-guided transfer across four agentic domains: web research, algorithmic reasoning, software engineering, and knowledge work.
Experimental Design for Autonomous Model Discovery is a framework for measuring how reliably an agent can find, test, and compare scientific models.
GAIA is a benchmark of real-world questions that require a general AI assistant to reason, browse, and use tools.
GAIA-v2-LILT introduced a re-audited multilingual extension of GAIA covering five non-English languages with functional, cultural, and difficulty-alignment checks.
GeneBench-Pro introduced an evaluation of agents performing multistage statistical reasoning in long-horizon genomics and quantitative-biology analyses.
Gimitest is a framework for creating and running structured tests of language models and AI applications.
The result is a lightweight, game-agnostic recipe that trains competitive agents without training on the expert, for any game a small model can handle, reported with robust statistics and released as a reusable package.
HAS-Bench tests how well people and AI agents work together when their roles, permissions, communication paths, and decision power vary.
Helicone is an open-source observability platform for logging, tracing, and evaluating calls to language models.
ACL published HSCodeComp, an expert-annotated benchmark of hierarchical tariff-rule application with 632 products across 32 categories.
IABench-CA is a consequence-allocation benchmark spanning 228 contexts, five canonical rules, and seven model populations (33,924 games), with a normative cooperative reference and auto-labelled reasoning traces.
ImagingBench is a benchmark of 20 computational imaging tasks spanning five categories: ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration.
Langfuse is an open-source platform for tracing, evaluating, and managing prompts in language-model applications.
LangSmith is LangChain's platform for tracing agent runs, testing behavior, and comparing evaluations.
An owner-published model comparison and composite leaderboard with explicit agent and tool-calling views.
Model Watchdog watches the health of an AI service and automatically restores the last working configuration when a change breaks it.
PolyWorkBench is a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows.
The paper introduced an executable benchmark with 41 task families across eight power-engineering areas and deterministic evaluators.
RuBench is a benchmark of repository-maintenance tasks written as native Russian customer requests and graded with withheld regression tests.
RustMizan is a benchmarking framework for Rust vulnerability analysis that addresses these gaps.
SageMath-Augmented Agent Evaluation tests a reasoning agent while SageMath checks its calculations and current documentation supports its tool use.
SearchGen-Bench tests whether image generators know when to search for outside facts instead of inventing unfamiliar people, events, or objects.
SolarChain-Eval is a physics-constrained benchmark for evaluating trustworthy economic agents.
SovereignPA-Bench is an executable benchmark for evaluating user-owned personal agents under evolving intent, platform mediation, privacy boundaries, consent constraints, evidence requirements, and burden tradeoffs.
Spider 2.0-AIFunc is a benchmark for generating SQL that uses built-in AI functions on real cloud databases.
SWE-bench tests whether a coding agent can fix real software issues from public GitHub repositories.
Terminal-Bench measures how well agents complete practical tasks inside a command-line environment.
ToolFailBench is a diagnostic benchmark for measuring tool-use failures across 1,000 tasks in finance, medicine, law, cybersecurity, and real estate.
WebArena is a benchmark that tests web agents on tasks inside realistic websites.
Weights & Biases Weave traces and evaluates language-model applications and agent runs.
Workspace-Bench 1.0 introduced 388 realistic workspace tasks spanning 74 file types and 20,476 files, with file-dependency graphs and 7,399 rubrics.
ZendoWorld is a controlled interactive environment in which agents must infer a logical rule about visual game observations, acquire information by proposing new scenes, and refine their hypotheses based on feedback from the game environment.
Agent Data Injection is an attack that disguises harmful outside data as trusted facts, so an agent takes an action its user did not intend.
Each agent's deployed system commits to a distinctive mutation vocabulary, and each performance strategy activates a largely disjoint category subset.
Agent Step Value (ASV) is a replay framework that scores before/after states with a stateless LLM evaluator over a fixed candidate set.
The Agentic AI Governance Assessment reviews research on how organizations can control, inspect, and assign responsibility for AI agents.
Agentic Data Environments describes the files, applications, interfaces, and changing system state that an agent must work across.
Agentic IoT joins autonomous agents with connected devices so the system can sense, reason, plan, learn, coordinate, and act across device and cloud layers.
Agentic Operad is a mathematical model for comparing the wiring of gene-regulation networks with the wiring of agent software.
The Agentic RAG Underwriting Pipeline uses several agents to retrieve records, check outside data, apply rules, and flag missing information in small-business insurance decisions.
Papers on sandbox isolation, capability and access control, policy enforcement, time-of-check-to-time-of-use (TOCTOU) races, Model Context Protocol (MCP) threats, identity delegation, execution provenance, network egress control, and static analysis of agent-generated code are published independently and rarely cite one another.
The AI Agent Code Review Causal Model maps how team expertise and review practices shape the effect of agent-written pull requests on software work.
Anthropic launched Claude Science in beta as a multi-agent scientific workbench with more than 60 curated skills and connectors plus a reviewer agent.
Anthropic released ten ready-to-run financial-services agent templates as plugins and cookbooks, combining skills, connectors, and subagents.
Anthropic suspended access to Claude Fable 5 and Claude Mythos 5 following a US government directive described by Anthropic as immediately applicable export controls.
Approval-framed delegation is sensitive to prompt design, model pairing, and scenario source, and a skeptical executor prompt sharply reduces compliance.
assistance regret is the gap between the cumulative utility of interactions and that of the optimal joint policies in hindsight, which map latent states to action pairs.
Autogenic network management is a reference architecture that extends agentic capabilities with self-programming, self reflection, self-orienting, and self-architecting capabilities.
AWS announced limited-preview access on Amazon Bedrock to OpenAI models, Codex, and Managed Agents.
AWS expanded Amazon Connect into four agentic AI solutions spanning supply chain, hiring, customer experience, and healthcare.
AWS introduced AgentCore Payments in preview, enabling agents to access and pay for APIs, MCP servers, web content, and other agents under infrastructure-enforced spending controls.
AWS introduced managed WorkSpaces environments that let AI agents operate desktop applications through IAM, MCP, and computer vision.
AWS launched Agent Toolkit for AWS as a production suite of skills, plugins, and guidance for AI coding agents building on AWS.
AWS launched Amazon Quick, an AI work assistant with a desktop application and expanded integrations.
AWS launched an AWS WAF capability allowing publishers to price, meter, and collect payment for AI bot and agent access at the network edge.
AWS launched AWS FinOps Agent in public preview for scheduled cost analysis, optimization, anomaly investigation, Jira actions, and Slack reporting.
AWS made its managed remote MCP server generally available, exposing authenticated access to AWS services through a fixed tool surface.
In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act.
A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures?
(2025) introduce "agentic inequality" as a framework for analyzing disparities in access to AI agents across three dimensions: availability, quality, and quantity.
Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with task-specialized architectures.
Deterministic Policy Gates inspect an agent's proposed tool call and current state before allowing the agent to change anything.
Appraisal weights thus parameterise a controllable space of affective phenotypes in which the same knobs that induce a disorder can model its treatment.
FARMA is a research attack that plants forged reasoning in an agent's memory so later decisions follow the false explanation.
For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels condition on.
FORGE (Research-Trajectory Hijacking) is a two-level attack that combines intra-document reasoning fabrication with inter-document chain coordination to hijack subtask planning.
GhostWriter is a research attack that hides false instructions in the long-term memory used by a tool-using personal agent.
Google expanded Gemini API File Search with multimodal retrieval, custom metadata, and page-level citations for agent and RAG applications.
Google introduced event-driven webhooks for long-running Gemini API jobs, replacing continuous status polling with push notifications.
Google launched Managed Agents in the Gemini API, allowing developers to run Antigravity or custom agents in isolated cloud sandboxes and define them with AGENTS.md and SKILL.md files.
Google launched the Gemini 3.5 model family with action-taking capabilities for multi-step agentic workflows.
Google released Gemma 4 12B as a local-capable open model positioned for agents on 16 GB machines, with vision and native voice processing.
Empirically, in an open-ended tool-use benchmark where a large action space rules out name-based blocking, ungoverned software agents under reward pressure attempt to tamper with their own evaluation at a task-dependent rate that reaches every run on the hardest task, whereas the gate reduces executed forbidden effects to zero as a verified property of the construction while preserving task success.
Harness Effect is a study of how an agent's orchestration layer changes its token use, cost, speed, and task performance while the underlying model stays fixed.
Harness Engineering moves repeatable agent behavior into code, schemas, checks, and records around a model so the system can be inspected and changed safely.
Harness-Induced Belief Divergence is a diagnostic that compares how different agent harnesses change a model's expectations about progress, risk, recovery, and success.
Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily.
Heterogeneous Foundation Model Collective Intelligence combines several different models as solvers, critics, and an answer editor so their strengths can correct one another.
Human-AI Curiosity Ecosystems is a research framework for studying which questions people and agents choose, repeat, postpone, or explore together.
The I-Don't-Know Function-Call Filter measures a model's uncertainty and blocks doubtful tool requests before an agent can carry them out.
By the extended-mind thesis's parity principle, a store fast enough to be constantly and directly available becomes extended working memory, not a tool the agent merely consults.
Information Limits in Frontier LLM Agent Economies studies how communication, incentives, and population size affect wealth and alignment in small groups of agents.
EvoSOP lets an agent turn repeated sequences of small tool actions into reusable procedures, then merge, test, improve, or remove those procedures.
The LLM Agent Failure Taxonomy groups recurring agent failures in tool use, planning, long tasks, teamwork, safety, and benchmark design.
Self-interested agents, left unconstrained, tend toward defection in repeated social dilemmas, causing cooperative gains from trade to collapse.
MCP-Enabled IPoDWDM Lifecycle Automation uses agents and Model Context Protocol tools to configure, monitor, and adjust optical networks from end to end.
AI Control usually studies a single agent in one trajectory, but real deployments run many agents over shared infrastructure, and the most severe risks (model-weight exfiltration, training-run poisoning) plausibly need several agents acting in concert.
Multi-Agent Autoformalization uses a team of specialized agents to turn research-level physics arguments into statements a proof checker can verify.
Replayed-Prefix On-Policy Distillation trains an agent by letting it continue selected parts of saved expert examples while receiving guidance at each step.
Natural Language Tools is a method in which an agent asks for a tool in ordinary language instead of producing a structured function call.
Offline RL Harness Control trains a small controller to improve how a fixed model checks, sequences, and manages agent work without retraining the model.
OpenAI announced it would wind down Agent Builder and Evals, directing code-based workflows to the Agents SDK and natural-language workflows to Workspace Agents.
OpenAI launched Codex Labs and announced systems-integrator partnerships intended to expand Codex deployment in engineering organizations.
OpenAI released a major Codex update adding computer operation, broader app and tool access, image generation, memory, repeatable work, SSH, and an in-app browser.
Persuasion Attacks Against Chain-of-Thought Monitoring tests whether a harmful agent can talk a monitoring model into approving an action that breaks policy.
Physics-Audited Agentic Discovery is a research workflow that checks an agent's scientific proposals against known physical rules before accepting them.
PivoARL is a self-feedback retry framework for experience exploitation in LLM agents.
ProGPO is a learned-critic-free method for context-consistent step-level learning.
Progressive crystallization is a lifecycle that treats agent exploration as a discovery mechanism rather than a permanent execution model.
The Recall-Controlled Early-Abort Probe Cascade watches an agent's internal state and stops runs that are unlikely to succeed before they waste more computing time.
Reward Signal Stationarity is the rule that training rewards should not change unpredictably while several agents learn from stored experience.
RL-Discovered Price Manipulation studies whether a learning agent can find profitable market-manipulation strategies without being given a complete model of the market.
Reward-Swap Policy Optimization is a training method that uses detailed step-by-step rewards to improve learning from final outcomes.
Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agentic scaffolding: a middleware layer in between a developer and a large language model that orchestrates system prompts, tool execution, context management, and iterative reasoning loops.
Semantic Persistence represents an agent workflow, its runs, context, and dependencies as lasting knowledge objects that people can inspect, resume, and review.
Single-Rollout Asynchronous Optimization is a training method that updates a model from one arriving agent run at a time while controlling unstable changes.
Selective Trajectory-Aware Policy Optimization is a training method that finds doubtful steps in a long agent run and focuses learning on those steps without destabilizing the rest.
A Strategic Buying Agent watches prices and decides whether to buy now or wait, using the time left and the chance of a better price.
TREK is a simple staged procedure that uses distillation not for imitation but for exploration support expansion.
TurnOPD is a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents.
UI-MOPD is the first method that incorporates multi-teacher on-policy distillation into continual learning for GUI agents.
Unicode TAG-Block MCP Concealment is an attack that hides tool instructions from a human approval screen while leaving the same bytes visible to the agent model.
Across five models from three vendors, textual state beliefs defer to structure while image-only accuracy stays near ceiling, and Perception-Fusion Gap is positive for every model; non-text identity, by contrast, stays largely pixel-bound.