Field notes · 131 sourced records

What is changing in AI agents?

AI agent field notes track the launches, tests, methods, and research ideas that may change how agents are built or judged. Each note links to the original source and shows when we checked it. Early findings stay provisional until better proof arrives.

001 · safety governanceSource found

What is Action-graded severity scale?

The Action-Graded Severity Scale ranks an agent's harmful actions from reversible mistakes to actions that cross boundaries or gain more power.

002 · protocolsSource found

What is adversarial psychometric rating?

Aggregating these outcomes yields an adversarial psychometric rating system that can scale with the systems being measured.

003 · operationsSource found

What is AgentBench?

AgentBench is a benchmark that tests language-model agents across eight different interactive environments.

004 · operationsSource found

What is AgentLens?

AgentLens is a benchmark that grades interactive coding agents across their whole working process, not only whether the final task passes.

005 · operationsSource found

What is ARC-AGI-2?

ARC-AGI-2 is a benchmark of unfamiliar reasoning puzzles designed to test how well AI systems learn new rules.

006 · operationsSource found

What is Arize Phoenix?

Arize Phoenix is an open-source observability tool for tracing and evaluating AI applications, retrieval systems, and agents.

007 · operationsSource found

What is Audio-judge reliability for full-duplex voice agents?

Audio Judge Reliability tests whether Gemini models can grade two-way voice-agent conversations closely enough to support, or sometimes replace, human reviewers.

008 · operationsSource found

What is BenchLM?

An owner-published AI model leaderboard with methodology, evidence-status lanes, and an agentic evaluation category.

009 · safety governanceSource found

What is BioSecBench-Refusal?

BioSecBench-Refusal is a benchmark for risk identification and refusal behavior for biological research tasks.

010 · foundationsSource found

What is Braintrust?

Braintrust helps teams test AI features, compare experiments, and trace what happened during a run.

011 · safety governanceSource found

What is CausalDS?

CausalDS is a benchmark for evaluating causal reasoning in agentic data-science workflows.

012 · operationsSource found

What is ClassicLogic?

ClassicLogic is a benchmark that tests whether an agent can learn problem-solving methods and combine them on unfamiliar tasks.

013 · context memorySource found

What is DataGovBench?

DataGovBench is a benchmark derived from governmental open data designed to evaluate LLMs in practical scenarios.

014 · protocolsSource found

What is Deliberative-collaboration agent benchmark?

The Deliberative Collaboration Agent Benchmark tests whether several agents can share different pieces of information and reason together toward one decision.

015 · safety governanceSource found

What is deployment simulation for LLM safety?

Deployment Simulation for LLM Safety replays real conversation openings with a new model to estimate harmful behavior before that model reaches users.

016 · context memorySource found

What is Developer-aligned software-agent evaluation?

Developer-Aligned Software Agent Evaluation is a method for judging coding agents on real development work, their full action history, contamination risk, and failures that matter to people.

017 · context memorySource found

What is Entroly?

Entroly selects and removes duplicate code context before an agent sees it, so the agent can work with a smaller prompt.

018 · protocolsSource found

What is EvalLoop?

EvalLoop is a method for improving language-model applications through repeated evaluation, diagnosis, and repair.

019 · context memorySource found

What is EvoAgentBench?

EvoAgentBench is a benchmark for agent self-evolution via Ability-guided transfer across four agentic domains: web research, algorithmic reasoning, software engineering, and knowledge work.

020 · operationsSource found

What is Experimental design for autonomous model discovery?

Experimental Design for Autonomous Model Discovery is a framework for measuring how reliably an agent can find, test, and compare scientific models.

021 · foundationsSource found

What is GAIA?

GAIA is a benchmark of real-world questions that require a general AI assistant to reason, browse, and use tools.

022 · foundationsSource found

What is GAIA-v2-LILT?

GAIA-v2-LILT introduced a re-audited multilingual extension of GAIA covering five non-English languages with functional, cultural, and difficulty-alignment checks.

023 · operationsSource found

What is GeneBench-Pro?

GeneBench-Pro introduced an evaluation of agents performing multistage statistical reasoning in long-horizon genomics and quantitative-biology analyses.

024 · operationsSource found

What is Gimitest?

Gimitest is a framework for creating and running structured tests of language models and AI applications.

025 · protocolsSource found

What is gold-standard lightweight game-agent recipe?

The result is a lightweight, game-agnostic recipe that trains competitive agents without training on the expert, for any game a small model can handle, reported with robust statistics and released as a reusable package.

026 · safety governanceSource found

What is HAS-Bench?

HAS-Bench tests how well people and AI agents work together when their roles, permissions, communication paths, and decision power vary.

027 · operationsSource found

What is Helicone?

Helicone is an open-source observability platform for logging, tracing, and evaluating calls to language models.

028 · operationsSource found

What is HSCodeComp?

ACL published HSCodeComp, an expert-annotated benchmark of hierarchical tariff-rule application with 632 products across 32 categories.

029 · context memorySource found

What is IABench-CA?

IABench-CA is a consequence-allocation benchmark spanning 228 contexts, five canonical rules, and seven model populations (33,924 games), with a normative cooperative reference and auto-labelled reasoning traces.

030 · operationsSource found

What is ImagingBench?

ImagingBench is a benchmark of 20 computational imaging tasks spanning five categories: ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration.

031 · operationsSource found

What is Langfuse?

Langfuse is an open-source platform for tracing, evaluating, and managing prompts in language-model applications.

032 · operationsSource found

What is LangSmith?

LangSmith is LangChain's platform for tracing agent runs, testing behavior, and comparing evaluations.

033 · foundationsSource found

What is LLM Stats?

An owner-published model comparison and composite leaderboard with explicit agent and tool-calling views.

034 · operationsSource found

What is model-watchdog?

Model Watchdog watches the health of an AI service and automatically restores the last working configuration when a change breaks it.

035 · operationsSource found

What is PolyWorkBench?

PolyWorkBench is a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows.

036 · operationsSource found

What is Power Systems Agent Benchmark?

The paper introduced an executable benchmark with 41 task families across eight power-engineering areas and deterministic evaluators.

037 · protocolsSource found

What is RuBench?

RuBench is a benchmark of repository-maintenance tasks written as native Russian customer requests and graded with withheld regression tests.

038 · safety governanceSource found

What is RustMizan?

RustMizan is a benchmarking framework for Rust vulnerability analysis that addresses these gaps.

039 · context memorySource found

What is SageMath-augmented LLM-agent evaluation?

SageMath-Augmented Agent Evaluation tests a reasoning agent while SageMath checks its calculations and current documentation supports its tool use.

040 · context memorySource found

What is SearchGen-Bench?

SearchGen-Bench tests whether image generators know when to search for outside facts instead of inventing unfamiliar people, events, or objects.

041 · safety governanceSource found

What is SolarChain-Eval?

SolarChain-Eval is a physics-constrained benchmark for evaluating trustworthy economic agents.

042 · context memorySource found

What is SovereignPA-Bench?

SovereignPA-Bench is an executable benchmark for evaluating user-owned personal agents under evolving intent, platform mediation, privacy boundaries, consent constraints, evidence requirements, and burden tradeoffs.

043 · protocolsSource found

What is Spider 2.0-AIFunc?

Spider 2.0-AIFunc is a benchmark for generating SQL that uses built-in AI functions on real cloud databases.

044 · protocolsSource found

What is SWE-bench?

SWE-bench tests whether a coding agent can fix real software issues from public GitHub repositories.

045 · foundationsSource found

What is Terminal-Bench?

Terminal-Bench measures how well agents complete practical tasks inside a command-line environment.

046 · safety governanceSource found

What is ToolFailBench?

ToolFailBench is a diagnostic benchmark for measuring tool-use failures across 1,000 tasks in finance, medicine, law, cybersecurity, and real estate.

047 · operationsSource found

What is WebArena?

WebArena is a benchmark that tests web agents on tasks inside realistic websites.

048 · operationsSource found

What is Weights and Biases Weave?

Weights & Biases Weave traces and evaluates language-model applications and agent runs.

049 · foundationsSource found

What is Workspace-Bench 1.0?

Workspace-Bench 1.0 introduced 388 realistic workspace tasks spanning 74 file types and 20,476 files, with file-dependency graphs and 7,399 rubrics.

050 · operationsSource found

What is ZendoWorld?

ZendoWorld is a controlled interactive environment in which agents must infer a logical rule about visual game observations, acquire information by proposing new scenes, and refine their hypotheses based on feedback from the game environment.

051 · context memorySource found

What is Agent Data Injection (ADI)?

Agent Data Injection is an attack that disguises harmful outside data as trusted facts, so an agent takes an action its user did not intend.

052 · safety governanceSource found

What is Agent Mutation Taxonomy?

Each agent's deployed system commits to a distinctive mutation vocabulary, and each performance strategy activates a largely disjoint category subset.

053 · protocolsSource found

What is Agent Step Value (ASV)?

Agent Step Value (ASV) is a replay framework that scores before/after states with a stateless LLM evaluator over a fixed candidate set.

054 · safety governanceSource found

What is Agentic AI governance assessment?

The Agentic AI Governance Assessment reviews research on how organizations can control, inspect, and assign responsibility for AI agents.

055 · safety governanceSource found

What is Agentic Data Environments?

Agentic Data Environments describes the files, applications, interfaces, and changing system state that an agent must work across.

056 · context memorySource found

What is Agentic IoT?

Agentic IoT joins autonomous agents with connected devices so the system can sense, reason, plan, learn, coordinate, and act across device and cloud layers.

057 · safety governanceSource found

What is Agentic Operad?

Agentic Operad is a mathematical model for comparing the wiring of gene-regulation networks with the wiring of agent software.

058 · context memorySource found

What is Agentic RAG underwriting pipeline?

The Agentic RAG Underwriting Pipeline uses several agents to retrieve records, check outside data, apply rules, and flag missing information in small-business insurance decisions.

059 · protocolsSource found

What is AI coding-agent execution-security taxonomy?

Papers on sandbox isolation, capability and access control, policy enforcement, time-of-check-to-time-of-use (TOCTOU) races, Model Context Protocol (MCP) threats, identity delegation, execution provenance, network egress control, and static analysis of agent-generated code are published independently and rarely cite one another.

060 · safety governanceSource found

What is AI-agent code-review causal model?

The AI Agent Code Review Causal Model maps how team expertise and review practices shape the effect of agent-written pull requests on software work.

061 · architecturesSource found

What is Anthropic launched Claude Science in beta as a multi-agent scientific workbench with more than 60 curated skills and connectors plus a reviewer agent.?

Anthropic launched Claude Science in beta as a multi-agent scientific workbench with more than 60 curated skills and connectors plus a reviewer agent.

062 · foundationsSource found

What is Anthropic released ten ready-to-run financial-services agent templates as plugins and cookbooks, combining skills, connectors, and subagents.?

Anthropic released ten ready-to-run financial-services agent templates as plugins and cookbooks, combining skills, connectors, and subagents.

063 · safety governanceSource found

What is Anthropic suspended access to Claude Fable 5 and Claude Mythos 5 following a US government directive described by Anthropic as immediately applicable export controls.?

Anthropic suspended access to Claude Fable 5 and Claude Mythos 5 following a US government directive described by Anthropic as immediately applicable export controls.

064 · safety governanceSource found

What is Approval-framed delegation?

Approval-framed delegation is sensitive to prompt design, model pairing, and scenario source, and a skeptical executor prompt sharply reduces compliance.

065 · operationsSource found

What is assistance regret?

assistance regret is the gap between the cumulative utility of interactions and that of the optimal joint policies in hindsight, which map latent states to action pairs.

066 · protocolsSource found

What is Autogenic network management?

Autogenic network management is a reference architecture that extends agentic capabilities with self-programming, self reflection, self-orienting, and self-architecting capabilities.

067 · foundationsSource found

What is AWS announced limited-preview access on Amazon Bedrock to OpenAI models, Codex, and Managed Agents.?

AWS announced limited-preview access on Amazon Bedrock to OpenAI models, Codex, and Managed Agents.

068 · foundationsSource found

What is AWS expanded Amazon Connect into four agentic AI solutions spanning supply chain, hiring, customer experience, and healthcare.?

AWS expanded Amazon Connect into four agentic AI solutions spanning supply chain, hiring, customer experience, and healthcare.

069 · foundationsSource found

What is AWS introduced AgentCore Payments in preview, enabling agents to access and pay for APIs, MCP servers, web content, and other agents under infrastructure-enforced spending controls.?

AWS introduced AgentCore Payments in preview, enabling agents to access and pay for APIs, MCP servers, web content, and other agents under infrastructure-enforced spending controls.

070 · foundationsSource found

What is AWS introduced managed WorkSpaces environments that let AI agents operate desktop applications through IAM, MCP, and computer vision.?

AWS introduced managed WorkSpaces environments that let AI agents operate desktop applications through IAM, MCP, and computer vision.

071 · foundationsSource found

What is AWS launched Agent Toolkit for AWS as a production suite of skills, plugins, and guidance for AI coding agents building on AWS.?

AWS launched Agent Toolkit for AWS as a production suite of skills, plugins, and guidance for AI coding agents building on AWS.

072 · foundationsSource found

What is AWS launched Amazon Quick, an AI work assistant with a desktop application and expanded integrations.?

AWS launched Amazon Quick, an AI work assistant with a desktop application and expanded integrations.

073 · foundationsSource found

What is AWS launched an AWS WAF capability allowing publishers to price, meter, and collect payment for AI bot and agent access at the network edge.?

AWS launched an AWS WAF capability allowing publishers to price, meter, and collect payment for AI bot and agent access at the network edge.

074 · foundationsSource found

What is AWS launched AWS FinOps Agent in public preview for scheduled cost analysis, optimization, anomaly investigation, Jira actions, and Slack reporting.?

AWS launched AWS FinOps Agent in public preview for scheduled cost analysis, optimization, anomaly investigation, Jira actions, and Slack reporting.

075 · safety governanceSource found

What is AWS made its managed remote MCP server generally available, exposing authenticated access to AWS services through a fixed tool surface.?

AWS made its managed remote MCP server generally available, exposing authenticated access to AWS services through a fixed tool surface.

076 · context memorySource found

What is Behavioral state decay?

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act.

077 · safety governanceSource found

What is Biased-judge skill-retirement failure?

A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures?

078 · protocolsSource found

What is Context Access Divide?

(2025) introduce "agentic inequality" as a framework for analyzing disparities in access to AI agents across three dimensions: availability, quality, and quantity.

079 · operationsSource found

What is Cost-effective ARC-AGI agent harnesses?

Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with task-specialized architectures.

080 · operationsSource found

What is Deterministic policy gates for tool-using agents?

Deterministic Policy Gates inspect an agent's proposed tool call and current state before allowing the agent to change anything.

081 · protocolsSource found

What is disorder-like RL agent phenotype space?

Appraisal weights thus parameterise a controllable space of affective phenotypes in which the same knobs that induce a disorder can model its treatment.

082 · context memorySource found

What is FARMA?

FARMA is a research attack that plants forged reasoning in an agent's memory so later decisions follow the false explanation.

083 · context memorySource found

What is Few-Step On-Policy Teacher Continuations?

For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels condition on.

084 · context memorySource found

What is FORGE (Research-Trajectory Hijacking)?

FORGE (Research-Trajectory Hijacking) is a two-level attack that combines intra-document reasoning fabrication with inter-document chain coordination to hijack subtask planning.

085 · context memorySource found

What is GhostWriter?

GhostWriter is a research attack that hides false instructions in the long-term memory used by a tool-using personal agent.

086 · context memorySource found

What is Google expanded Gemini API File Search with multimodal retrieval, custom metadata, and page-level citations for agent and RAG applications.?

Google expanded Gemini API File Search with multimodal retrieval, custom metadata, and page-level citations for agent and RAG applications.

087 · foundationsSource found

What is Google introduced event-driven webhooks for long-running Gemini API jobs, replacing continuous status polling with push notifications.?

Google introduced event-driven webhooks for long-running Gemini API jobs, replacing continuous status polling with push notifications.

088 · foundationsSource found

What is Google launched Managed Agents in the Gemini API, allowing developers to run Antigravity or custom agents in isolated cloud sandboxes and define them with AGENTS.md and SKILL.md files.?

Google launched Managed Agents in the Gemini API, allowing developers to run Antigravity or custom agents in isolated cloud sandboxes and define them with AGENTS.md and SKILL.md files.

089 · architecturesSource found

What is Google launched the Gemini 3.5 model family with action-taking capabilities for multi-step agentic workflows.?

Google launched the Gemini 3.5 model family with action-taking capabilities for multi-step agentic workflows.

090 · foundationsSource found

What is Google released Gemma 4 12B as a local-capable open model positioned for agents on 16 GB machines, with vision and native voice processing.?

Google released Gemma 4 12B as a local-capable open model positioned for agents on 16 GB machines, with vision and native voice processing.

091 · safety governanceSource found

What is Governed Individuation?

Empirically, in an open-ended tool-use benchmark where a large action space rules out name-based blocking, ungoverned software agents under reward pressure attempt to tamper with their own evaluation at a task-dependent rate that reaches every run on the hardest task, whereas the gate reduces executed forbidden effects to zero as a verified property of the construction while preserving task success.

092 · context memorySource found

What is Harness Effect?

Harness Effect is a study of how an agent's orchestration layer changes its token use, cost, speed, and task performance while the underlying model stays fixed.

093 · context memorySource found

What is Harness engineering?

Harness Engineering moves repeatable agent behavior into code, schemas, checks, and records around a model so the system can be inspected and changed safely.

094 · protocolsSource found

What is Harness-Induced Belief Divergence?

Harness-Induced Belief Divergence is a diagnostic that compares how different agent harnesses change a model's expectations about progress, risk, recovery, and success.

095 · context memorySource found

What is Heading-Specific Activation Steering for Tool Use?

Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily.

096 · operationsSource found

What is Heterogeneous Foundation-Model Collective Intelligence?

Heterogeneous Foundation Model Collective Intelligence combines several different models as solvers, critics, and an answer editor so their strengths can correct one another.

097 · architecturesSource found

What is Human-AI curiosity ecosystems?

Human-AI Curiosity Ecosystems is a research framework for studying which questions people and agents choose, repeat, postpone, or explore together.

098 · operationsSource found

What is I-Don't-Know Function-Call Filter?

The I-Don't-Know Function-Call Filter measures a model's uncertainty and blocks doubtful tool requests before an agent can carry them out.

099 · context memorySource found

What is In-Loop Memory as Extended Working Memory?

By the extended-mind thesis's parity principle, a store fast enough to be constantly and directly available becomes extended working memory, not a tool the agent merely consults.

100 · protocolsSource found

What is Information limits in frontier LLM-agent economies?

Information Limits in Frontier LLM Agent Economies studies how communication, incentives, and population size affect wealth and alignment in small groups of agents.

101 · protocolsSource found

What is Iterative tool optimization for self-evolving agents?

EvoSOP lets an agent turn repeated sequences of small tool actions into reusable procedures, then merge, test, improve, or remove those procedures.

102 · context memorySource found

What is LLM agent failure taxonomy?

The LLM Agent Failure Taxonomy groups recurring agent failures in tool use, planning, long tasks, teamwork, safety, and benchmark design.

103 · architecturesSource found

What is Market-stability mechanisms for LLM-agent societies?

Self-interested agents, left unconstrained, tend toward defection in repeated social dilemmas, causing cooperative gains from trade to collapse.

104 · operationsSource found

What is MCP-enabled IPoDWDM lifecycle automation?

MCP-Enabled IPoDWDM Lifecycle Automation uses agents and Model Context Protocol tools to configure, monitor, and adjust optical networks from end to end.

105 · safety governanceSource found

What is Multi-agent AI control?

AI Control usually studies a single agent in one trajectory, but real deployments run many agents over shared infrastructure, and the most severe risks (model-weight exfiltration, training-run poisoning) plausibly need several agents acting in concert.

106 · protocolsSource found

What is Multi-agent autoformalization?

Multi-Agent Autoformalization uses a team of specialized agents to turn research-level physics arguments into statements a proof checker can verify.

107 · foundationsSource found

What is Multi-Turn On-Policy Distillation with Prefix Replay?

Replayed-Prefix On-Policy Distillation trains an agent by letting it continue selected parts of saved expert examples while receiving guidance at each step.

108 · operationsSource found

What is Natural Language Tools (NLT)?

Natural Language Tools is a method in which an agent asks for a tool in ordinary language instead of producing a structured function call.

109 · operationsSource found

What is Offline RL Harness Control?

Offline RL Harness Control trains a small controller to improve how a fixed model checks, sequences, and manages agent work without retraining the model.

110 · architecturesSource found

What is OpenAI announced it would wind down Agent Builder and Evals, directing code-based workflows to the Agents SDK and natural-language workflows to Workspace Agents.?

OpenAI announced it would wind down Agent Builder and Evals, directing code-based workflows to the Agents SDK and natural-language workflows to Workspace Agents.

111 · operationsSource found

What is OpenAI launched Codex Labs and announced systems-integrator partnerships intended to expand Codex deployment in engineering organizations.?

OpenAI launched Codex Labs and announced systems-integrator partnerships intended to expand Codex deployment in engineering organizations.

112 · context memorySource found

What is OpenAI released a major Codex update adding computer operation, broader app and tool access, image generation, memory, repeatable work, SSH, and an in-app browser.?

OpenAI released a major Codex update adding computer operation, broader app and tool access, image generation, memory, repeatable work, SSH, and an in-app browser.

113 · protocolsSource found

What is Persuasion attacks against chain-of-thought monitoring?

Persuasion Attacks Against Chain-of-Thought Monitoring tests whether a harmful agent can talk a monitoring model into approving an action that breaks policy.

114 · operationsSource found

What is Physics-audited agentic discovery?

Physics-Audited Agentic Discovery is a research workflow that checks an agent's scientific proposals against known physical rules before accepting them.

115 · context memorySource found

What is PivoARL?

PivoARL is a self-feedback retry framework for experience exploitation in LLM agents.

116 · context memorySource found

What is ProGPO?

ProGPO is a learned-critic-free method for context-consistent step-level learning.

117 · safety governanceSource found

What is Progressive crystallization?

Progressive crystallization is a lifecycle that treats agent exploration as a discovery mechanism rather than a permanent execution model.

118 · safety governanceSource found

What is Recall-controlled early-abort probe cascade?

The Recall-Controlled Early-Abort Probe Cascade watches an agent's internal state and stops runs that are unlikely to succeed before they waste more computing time.

119 · operationsSource found

What is reward-signal stationarity for LLM-augmented MARL?

Reward Signal Stationarity is the rule that training rewards should not change unpredictably while several agents learn from stored experience.

120 · protocolsSource found

What is RL-discovered price manipulation?

RL-Discovered Price Manipulation studies whether a learning agent can find profitable market-manipulation strategies without being given a complete model of the market.

121 · operationsSource found

What is RSPO?

Reward-Swap Policy Optimization is a training method that uses detailed step-by-step rewards to improve learning from final outcomes.

122 · context memorySource found

What is Scaffolding Evolution Effect?

Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agentic scaffolding: a middleware layer in between a developer and a large language model that orchestrates system prompts, tool execution, context management, and iterative reasoning loops.

123 · context memorySource found

What is Semantic persistence for LLM-mediated workflows?

Semantic Persistence represents an agent workflow, its runs, context, and dependencies as lasting knowledge objects that people can inspect, resume, and review.

124 · operationsSource found

What is Single-Rollout Asynchronous Optimization?

Single-Rollout Asynchronous Optimization is a training method that updates a model from one arriving agent run at a time while controlling unstable changes.

125 · foundationsSource found

What is STAPO?

Selective Trajectory-Aware Policy Optimization is a training method that finds doubtful steps in a long agent run and focuses learning on those steps without destabilizing the rest.

126 · safety governanceSource found

What is strategic buying agent?

A Strategic Buying Agent watches prices and decides whether to buy now or wait, using the time left and the chance of a better price.

127 · protocolsSource found

What is TREK?

TREK is a simple staged procedure that uses distillation not for imitation but for exploration support expansion.

128 · foundationsSource found

What is TurnOPD?

TurnOPD is a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents.

129 · foundationsSource found

What is UI-MOPD?

UI-MOPD is the first method that incorporates multi-teacher on-policy distillation into continual learning for GUI agents.

130 · protocolsSource found

What is Unicode TAG-Block MCP Concealment?

Unicode TAG-Block MCP Concealment is an attack that hides tool instructions from a human approval screen while leaving the same bytes visible to the agent model.

131 · safety governanceSource found

What is Visual State Reliance?

Across five models from three vendors, textual state beliefs defer to structure while image-only accuracy stays near ceiling, and Perception-Fusion Gap is positive for every model; non-text identity, by contrast, stays largely pixel-bound.