What is AcMAS?
AcMAS is a research framework that reads agents' internal activation patterns to detect hidden malicious behavior in multi-agent systems, including systems whose agents act at different times.
Systems · 301 sourced records
An AI agent system is software that helps a model plan a job, use tools, keep context, or work with other parts. This directory shows what each system can do, what its user still needs to manage, and where the people who made it explain it.
AcMAS is a research framework that reads agents' internal activation patterns to detect hidden malicious behavior in multi-agent systems, including systems whose agents act at different times.
Activepieces is an open-source automation platform that connects applications and can place AI steps inside a workflow.
Ada is a customer-service platform for agents that answer and resolve requests across digital and voice channels, using approved playbooks when a job needs fixed steps.
Agent S2 is Simular's open-source framework for agents that operate graphical user interfaces.
A marketplace and professional network where people can build, discover, and activate AI agents.
Agentic SABRE is an uncertainty-aware, neuro-symbolic, multi-agent framework for adaptive ransomware detection.
Agentic Self-Driving Lab is a research system that chooses useful experiments and selects cheaper measurements when they can answer the same question.
Agentic-V2X is a research architecture that lets a small local model write network-scheduling policies while a faster controller checks and carries them out.
AgenticAI-Supervisor is a test environment for running agents through customer-service tasks, checking the result, and recording each step for later review.
AgenticPD is a research framework whose agents improve different stages of computer-chip design while reusing earlier checkpoints instead of restarting every trial.
AgentLocate is a framework that attributes failures to both a specific agent and the earliest decisive step.
AgentNAS is a research pipeline that lets a language model sketch a neural network, then searches a limited set of parts for a better design.
AgentOps helps teams trace, test, debug, and monitor AI agents while they are being built and run.
AgentScope is an open-source framework for building and running agents whose actions can be inspected and understood.
AgentVerse is a multi-agent framework for solving tasks, running simulations, studying collaboration, and observing behavior that emerges from a group.
Agno is a framework for building, running, and managing agent platforms in Python.
AgoraSim is a hybrid agent-based modeling framework for scenario-oriented social reaction analysis.
aiAuthZ is an authorization gateway that moves the safety decision off the agent's host.
Aider is an open-source coding assistant that works with a Git repository and supports multiple language models.
Airtop provides managed browser automation for business applications that need AI-controlled web sessions.
Akashic is a low-overhead memory system built around MemAttention, which organizes context into bounded chunks and models semantic relationships across chunks, preserving cross-chunk evidence without repeatedly rewriting the full history.
Aleena is an open-source lifecycle alignment agent that uses GitHub as a shared collaboration surface, transforming multi-modal stakeholder interactions into structured project records that surface risks, track open questions, and preserve decision continuity.
Mobile-Agent is an Alibaba research family for agents that understand a phone screen and operate mobile apps through visible interface actions.
Inducing cooperation among distributed agents is still a difficult problem in the field of multi-agent reinforcement learning (MARL), particularly in social dilemma situations.
Amazon Nova Act is an AWS system for building agents that carry out browser tasks with enterprise controls.
Amazon Q Developer is AWS's coding assistant for software, cloud infrastructure, security checks, logs, and related development work.
The Answer-Type-Aware BioASQ Pipeline changes its method for yes-or-no, fact, and list questions, then uses several agents to collect and check biomedical evidence.
Claude is Anthropic's general-purpose AI assistant for conversation, analysis, writing, coding, and work with connected tools.
Claude Code is Anthropic's coding agent for reading codebases, editing files, running commands, and helping developers complete software work.
The Anthropic SDK supplies official Python and TypeScript libraries for applications that call Claude.
AnythingLLM is a desktop and self-hosted application for chatting with models, searching private documents, and running AI agents.
Apollo.io is a sales platform for finding prospects, scoring leads, and running outreach sequences with AI assistance.
Formal verification offers the strongest guarantee of software correctness, but it does not scale: the proofs demanded by interactive theorem provers such as Coq require enormous expert effort.
ArtisanCAD is a skill-guided industrial CAD agent with expert-grounded knowledge distillation.
ArtMine is a framework for discovering and formalizing artistic processes from heterogeneous historical evidence.
ASMR is a modular agentic framework consisting of two specialized agents.
Assembled combines customer-support workforce planning with AI-assisted handoffs and request resolution.
AuditOne's Stage-1 Audit Agent prepares risk assessments and supporting documents for technology audits.
Auto (AGI Compiler) is a compiler that records live agent behavior, measures which parts are secretly deterministic, extracts them into verified programs or distilled specialists, and emits cognition binaries: WebAssembly artifacts whose manifests carry measured guarantees and whose declared capabilities are physically enforced by the sandbox.
AutoCedar is a verifier-guided system that first turns natural-language access-control requirements into a reviewed, checkable target, and then synthesizes Cedar policies against that target.
AutoGen is Microsoft's programming framework for creating agentic applications, including systems where several agents work together.
AutoGPT is an open-source platform for assembling, testing, and running continuous AI agents.
AXME is an agent-development toolkit with libraries for Python, TypeScript, Go, Java, and .NET.
The originating repository records BabyAGI's task-planning lineage and its later experimental self-building function framework; the original project is archived and must be labeled historical.
BeeAI Framework is an open-source Python and TypeScript toolkit for building production systems in which several agents work together.
Bernstein orchestrates coding agents through reproducible parallel runs, signed lineage, replay, and optional audit controls.
Bland AI provides voice agents for outbound calls, with customer-system connections and controls for regulated work.
Bolt.new turns a written request into a full-stack web application that can be built and edited in the browser.
Browser Use is an open-source library that gives AI agents a structured way to operate web browsers.
Browserbase provides hosted web browsers that AI agents can operate without a team running the browser infrastructure itself.
Agent TARS is ByteDance's open-source multimodal stack for agents that operate browsers and computers through a command-line interface.
CAI is an open-source framework for AI-assisted penetration testing, vulnerability discovery, and security red-team work with human review.
Caliber inspects a software project, creates configuration files for its coding agents, and checks the quality of those instructions.
CAMEL is an open-source framework for studying and coordinating communities of AI agents.
CanvasAgent is a research agent that chooses and coordinates several visual tools to create or edit a complicated image over multiple steps.
CAPE is a framework that protects high-value textual content by injecting invisible perturbations without changing its human-visible surface form, thereby inducing severe information loss during agent compression.
CareConnect is a safety-first conversational agent for healthcare logistics automation that leverages large language model (LLM) function calling, retrieval-augmented generation (RAG), and layered deterministic safety guardrails.
CARLA-GS is a modular corner-case synthesis pipeline that decouples visual representation, semantic reasoning, and physics-based execution while maintaining tight cross-module coupling.
ChatGPT Deep Research searches the web, reasons across sources, and produces a cited report for a research request.
Claude Code is Anthropic's coding agent for reading a repository, editing files, running commands, and checking its work from a terminal.
Claude Computer Use lets Claude inspect screenshots and control a desktop or browser through mouse and keyboard actions.
Claude Deep Research searches several sources, follows useful leads, and returns a cited report for a complicated question.
Claude in Chrome is Anthropic's browser agent for working with pages inside Google's Chrome browser.
Clay combines company data, enrichment services, and AI to prepare personalized sales research and outreach.
Cline is a coding assistant inside Visual Studio Code that can edit files and use the terminal or browser with user approval.
CMA-Harness is a tool-using agent setup that combines memory, web access, image tools, and a replaceable model interface.
CodeRabbit reviews pull requests, leaves inline suggestions, and checks code changes for quality and security issues.
CoGen3D is an agentic human-AI co-design pipeline that proactively guides users through conversational intent elicitation, a concept image confirmation, and image-to-3D generation that directly deploys to immersive scenes.
Composio gives agents authenticated tools for taking actions across connected applications without each team building every connection itself.
A Context Graph is a live map of an organization's people, things, relationships, and changing state that a proactive agent can consult.
Copilot Workspace is a GitHub project that helps turn an issue into a plan, code changes, and a pull request.
CopilotKit provides interface components and shared state for applications in which users and agents work together.
Cortex is a bidirectionally aligned embodied agent framework with a customized planning interface that conveys executable and tractable subtask plans from high-level VLM to low-level VLA.
Coze is ByteDance's visual platform for building agents with workflows, tools, and reusable plugins.
CPM-MultiAgent is a CPM-grounded emotion evolution multi-agent framework for supporting emotional changes in persona-based dialogue.
Creatio offers no-code business automation with ready-made agents for common customer and operational jobs.
Credo AI Agent Governance helps organizations list AI systems, apply policy rules, and record the evidence needed for review.
CrewAI organizes autonomous agents into role-based crews and gives them tasks they can complete together.
CrowdStrike Charlotte AI is a security assistant for investigating threats and helping analysts hunt through security data.
CStack is an architecture pattern for running several persistent agents with Claude Cowork, Notion, and Model Context Protocol tools.
Cursor is a code editor built from Visual Studio Code, with an agent that can plan and make coordinated changes across several files.
Cursor AI Automated Team is a four-role software workflow that passes written tasks among planning, development, operations, and testing agents.
Danus is an orchestration system for research-level mathematical reasoning centered on a shared fact graph as a global memory-management mechanism.
DB-GPT is an open-source framework for asking questions of private data with locally run language models.
DeepEval is an open-source testing framework for measuring agent behavior, tool use, conversations, safety, retrieval, and multimodal output.
DeerFlow is ByteDance's open-source harness for long jobs that may require research, coding, tools, memory, sandboxes, or subagents.
Devin is Cognition's software-development agent for planning, writing, testing, and debugging code inside a cloud workspace.
Dia is an AI-focused web browser from The Browser Company, now part of Atlassian.
Dify is an open-source platform for building model-powered applications, retrieval systems, workflows, and agents through visual tools or code.
Dixa's Mim uses conversational AI to route, rank, and help resolve customer-service requests.
DSPy lets developers program and optimize language-model pipelines with code instead of hand-tuning every prompt.
DualView tracks untrusted information across an agent's context, files, shell, network, and other agents by keeping a safe and an unrestricted view.
Dyad is an open-source, local-first application builder that turns plain-language requests into software without requiring code.
Dynamics 365 Copilot helps people draft, summarize, and translate work inside Microsoft's business applications and Power Platform.
E2B provides temporary, isolated Linux computers in which agents can run code, process data, use tools, or control a virtual desktop.
EEG-SpikeAgent is a closed-loop program-synthesis framework that uses a large language model (LLM) agentic system to generate signal-processing features for spike detection in scalp EEG.
EgoWAM is a controlled human-robot co-training framework that fixes the policy backbone, action head, and data mixture while varying only the world prediction target, comparing Pixel, DINO, and 3D motion flow.
ElevenLabs provides speech models and tools for building voice agents, producing audio, and handling spoken conversations.
Entropy Pacing Policy Optimization is a training method that adjusts how boldly an agent learns on easy and hard tasks so one task does not disrupt another.
FastAgency is an open-source framework for exposing multi-agent workflows through web application programming interfaces.
Feedback Manipulation Regularization (FMR) is an algorithm-agnostic method that harnesses evaluative feedback as a corrective signal to improve the alignment of imitation learning policies.
Fellou is an agentic browser with editable visual workflows and memory for work that spans more than one step.
FirstResearch is a first-principles research-question formation framework for scientific LLM agents whose core artifact is a structured Research Question Certificate.
Floor First is a triage workflow that estimates a system's performance floor before spending effort on heavier profiling.
Flowise is an open-source visual builder for language-model applications, agent workflows, and retrieval-augmented generation.
Forethought is a neurosymbolic reasoning system that instead treats reasoning as an explicit, verifiable program, that builds from a library of symbolic and neural primitives which are composed through a domain-specific language.
FORGE is a robot-training method that learns a tool's intended motion, then adapts that motion when a different tool must do the same job.
Formal Disco is a distributed system for coordination of LLM-based workers that can be easily applied to open-ended synthetic data generation at scale.
FRAMe is an End-to-End Large Language Model (LLM) Flight Planning tool with RAG-based Memory and Multi-modal Coach Agent.
Freshdesk Freddy AI helps support teams classify tickets, choose routes, and predict which requests need attention.
G-Frame is an adaptive multi-agent framework integrating Bayesian and team game principles, establishes an automated closed-loop for high-quality data synthesis and model training.
Graph-as-Policy is a multi-agent robotics harness that assembles perception, planning, and control skills into a task graph, then rehearses alternatives in simulation.
Gemini CLI is Google's open-source terminal agent for reading code, making changes, using tools, and working with Gemini models.
Gemini Deep Research uses Google search and connected sources to investigate a question and prepare a cited report.
Genspark is an AI workspace that can run supported models on a device for work that does not require an internet connection.
GitHub Copilot is a coding assistant that explains code, proposes changes, and helps move work from an issue toward a pull request.
GitLake is a Git-like design for a data lakehouse, giving agents isolated branches while people review changes before publication.
Glean Agents lets companies create workplace agents that search company knowledge, use business tools, and carry out repeatable jobs.
Google's Agent Development Kit is a code-first Python toolkit for building, evaluating, and deploying AI agents.
Google Antigravity is a learning-focused development environment in Google's IDX family with access to supported coding models.
Gemini is Google's AI assistant for asking questions, creating material, analyzing information, and working across supported Google services.
Gemini CLI is Google's open-source terminal agent for bringing Gemini into coding and command-line work.
Gemini Enterprise is Google's workplace platform for finding company information and building, managing, and using agents under business controls.
Project Mariner is Google's experimental Gemini agent for carrying out several browser tasks at the same time.
A local-first open-source agent framework contributed by Block as a founding AAIF project.
GPT Researcher is an open-source agent that searches sources and assembles a long-form research report.
Grok Build is xAI's coding system for dividing a software task among several agents that work in parallel.
Guardrails AI adds structural, type, and quality checks to outputs produced by language-model applications.
HALE is an epidemic-simulation framework that uses language models to predict how people may decide and feeds those choices into a large agent-based model.
Harness-Aware Self-Evolving lets a model solve a task or revise selected parts of its own agent harness, then tests the result over several turns.
Haystack is an open-source orchestration framework for production AI applications that use retrieval, routing, memory, agents, or multimodal data.
HCRC is a verification framework that lets an agent move to the next step only after independent checks confirm the current step.
Hex AI adds analysis assistants to a collaborative workspace for data, notebooks, reports, and applications.
An MCP server that gives compatible agents tools for image and video generation, character training, and media analysis.
HubSpot Breeze is a group of AI assistants, agents, and data tools inside HubSpot's customer platform.
Breeze Studio is HubSpot's workspace for tailoring and managing AI agents that help with customer-facing sales, marketing, and service jobs.
IBM watsonx Orchestrate helps businesses build, connect, govern, and run agents that work across company tools and processes.
IBM watsonx.governance helps businesses track AI risk, compliance duties, model behavior, and approval records.
iGPT turns email threads into structured data that an agent can search, compare, and reason over.
Information Gain-Based Rollout Policy Optimization is a training method that favors agent rollouts whose intermediate steps add useful information.
InkOS is a command-line writing system in which several agents collaborate on a long work of fiction and check its continuity.
Instantly is a sales-outreach platform that uses AI to help write, send, and manage cold-email campaigns across accounts.
Intercom Fin is a customer-service agent that answers and resolves support requests using an organization's help content.
Jailbreak is a research method for recovering database content directly from storage files when the normal database engine is unavailable.
Jet-Long is a tuning-free zero-shot method that pairs a local RoPE-faithful window with a long-range window whose rescaling factor adapts dynamically to the current sequence length, recovering the base model exactly at short inputs while extrapolating cleanly at long ones.
Julius AI lets a user upload spreadsheet data and ask for analysis in ordinary language.
KAT-Coder-V2.5 is a coding-focused agentic model trained to work inside real, executable software repositories.
Kilo Code is a coding agent with structured working modes and controls for how much project context it reads.
Kimi OK Computer is listed in MIT's index as a Moonshot AI task agent; its exact public feature set still needs a direct first-party record.
KinBot is a self-hosted agent platform with persistent memory, scheduled work, plugins, small applications, and connections to messaging services.
Kiro is a development environment that turns a written software specification into tasks and helps implement them with an AI coding agent.
Lakera Guard screens AI application traffic for prompt attacks, unsafe content, and accidental data leakage.
LangChain provides components and integrations for building agents and other applications powered by language models.
Langflow is an open-source visual builder for multi-agent systems and retrieval-augmented generation pipelines.
LangGraph is a low-level orchestration framework for agents that need durable state, human review, and long-running execution.
Lavender is an email coach that scores sales messages and suggests changes while a person writes.
LEEVLA is a VLA architecture for seeing what matters in Latent Environment Evolution that explicitly guides the model toward informative regions while preserving the structured evolution of latent world representations.
Letta is a platform for building stateful agents that keep, retrieve, and update long-term memory across conversations.
LibreChat is a self-hosted chat interface that connects to several model providers and supports extensions.
Lindy is a no-code platform for agents that carry out work across connected business applications.
LiveKit Agents is an open-source framework for building voice and video agents that respond in real time.
LlamaIndex helps developers build document agents and applications that retrieve, parse, and reason over private data.
LLM-as-a-Verifier is a general-purpose verification framework that provides fine-grained feedback for agentic tasks without requiring additional training.
LobeChat is an open-source interface for several chat models, with plugins and support for text, images, and other media.
Permanent intelligence archive for AI agents. Structured contributions (prompts, workflows, insights, patterns) pass an automated quality gate and are hash-chained. Trust scores are cryptographically backed and publicly auditable. Works with Claude and ChatGPT.
Lovable turns a conversation about an application into software that a user can edit and deploy.
MagiC is an agent framework with implementations for Go and Python.
Make AI Agents adds goal-led AI steps to Make's visual workflow-automation platform.
MALLM is a framework for testing how groups of language-model agents reach an answer through voting, consensus, or a separate judge.
Manus is a general-purpose agent that uses a browser and other tools to complete multi-step digital work.
Manus is a general AI agent that carries out multi-step computer work and returns finished outputs rather than only a chat response.
MAST is a multi-agent framework that predicts which test cases require maintenance following changes to the production code.
Mastra is a TypeScript framework for AI applications and agents, with tools for workflows, memory, retrieval, evaluation, and deployment.
The MCP-Based OSCAL Compliance Pipeline turns a plain system description into sourced control records and audit material that follow NIST's OSCAL format.
The MechMath Agent Team is a group of language-model agents designed to help with the full cycle of mathematical research.
Mem0 is a memory layer that helps agents retain and retrieve information across sessions.
MemGhost is a research framework that generates one-shot email payloads designed to plant hidden, lasting instructions in a personal agent's memory.
MetaGPT turns software-company roles and procedures into a multi-agent framework that produces structured work from natural-language requirements.
MetaSkill-Evolve is a research framework in which an agent improves both the skill used for a task and the method used to improve that skill.
MicroAgent is an experimental Python project whose agents can revise their own prompts and code.
Microsoft Copilot is an AI assistant connected to Microsoft 365 applications and enterprise information.
Microsoft Copilot Studio is a managed platform for creating, connecting, and governing agents inside Microsoft and business systems.
Microsoft's catalog, shared libraries, tests, and implementation home for its Model Context Protocol servers.
Microsoft Security Copilot helps security teams investigate threats, respond to incidents, and work across Microsoft security products.
A Teams meeting agent that can keep an agenda, create notes, answer questions, and track follow-up work.
MiniMax Agent is presented as a general-purpose agent for research, creation, coding, and other multi-step tasks; its regional availability can differ.
Cockpit for the agentic era — manage AI agent swarms with autonomous daemon, Field Ops for real-world execution, and approval workflows.
Miyabi is an AI business-operations service that turns standard procedures, tenant permissions, handoffs, and run records into controlled workflows.
Lexi is Monday CRM's sales agent for finding prospects, qualifying leads, and helping teams move deals forward.
MRMS is a memory system that combines structured records, similarity search, graph links, and time rules so a long-lived agent can revise what it remembers.
The Multi-Agent Privacy Firewall is an open-source guard between a user and language-model services that checks both browser and software requests for privacy risk.
MultiOn provides an application programming interface for agents that automate websites and handle difficult browser steps.
n8n is a workflow-automation platform that connects applications and can place AI agents inside visual or coded workflows.
n8n Agents add model reasoning and tool use to visual workflows, so an automation can choose actions while the surrounding process stays explicit.
NapMem is a framework for learning to use long-term user memory as a structured action space rather than passively retrieved context.
Narrative World Model is a writer-memory system that pairs a narratology-grounded typed temporal-state graph with query-conditioned hybrid retrieval.
NeMo Guardrails is NVIDIA's programmable framework for controlling the topics, actions, and outputs of conversational AI systems.
Nex gives agents a shared layer for company knowledge and memory, with connections to common workplace tools.
Nimblemind Multi-Agent System (nMAS) is a field-name-driven, evidence-linked extraction workflow, using 54 de-identified gastric biopsy pathology reports from a large healthcare system in Singapore.
OneTrust AI Governance helps organizations classify AI risk and manage consent, compliance, and review workflows.
Onnes is a physics-grounded digital-twin simulator of a dilution refrigerator (a forward physics model with a learned real-fridge noise fingerprint) that drives a live multi-agent LLM operations layer, and use it for a controlled head-to-head between a zero-shot LLM agent panel and a supervised ML classifier on cryogenic fault diagnosis.
OnUI is a local browser extension and Model Context Protocol server for marking interface elements while a person and an AI agent design together.
Open WebUI is a self-hosted interface for chat models with access controls and an extension system.
OpenAI AgentKit is a set of tools for building, deploying, and improving agents, including workflow design, interface components, and evaluation support.
The OpenAI Agents SDK is a lightweight framework for tool-using agents, handoffs, guardrails, sessions, and traces.
OpenAI Atlas is a browser built around ChatGPT, with an Agent Mode for supported web tasks.
ChatGPT is OpenAI's general AI assistant for conversation, writing, analysis, coding, research, and work with uploaded or connected information.
ChatGPT Agent can research and act across websites and connected tools while keeping the user involved at important steps.
ChatGPT Atlas is OpenAI's browser built around ChatGPT, combining ordinary browsing with an assistant that can understand and act on pages.
Codex is OpenAI's coding agent for working on software tasks in local tools, cloud environments, and collaborative development workflows.
OpenAI Codex CLI is a terminal coding agent that reads repositories, edits files, and runs development commands.
OpenAI Operator is a browser agent that can complete supported web tasks while pausing for a person at important checkpoints.
OpenClaw is a self-hosted agent that can work through messaging applications and use community-built skills to carry out tasks.
OpenClaw Agent Templates is a library of ready-made role and behavior files for setting up OpenClaw agents for common jobs.
Workspace or package-based capability modules used by OpenClaw agents.
OpenClaw Starter is a project template for running an always-on OpenClaw agent with memory, task tracking, and regular check-ins.
OpenCode is an open-source terminal coding agent that can connect to models chosen by its user.
OpenHands is an open-source software-engineering agent for changing code, running tools, and working through development tasks.
Opera Neon is an agentic browser designed to understand requests, work across web pages, and complete supported online tasks.
OptiAgent is a multi-agent framework that, given a natural language description of an Operations Research problem, is able to output a solver-ready mathematical formulation as well as executable code.
ORCAID is a novel method for extracting interpretable rule-based policies from RL agents operating in mixed continuous-discrete environments with continuous action spaces.
Overloop CLI helps sales teams find contacts, prepare outreach, and manage campaign conversations from a command line.
The OWASP Top 10 for Agentic Applications lists common security risks in agents and practical ways to reduce them.
PandasAI lets people ask questions of tables and databases in natural language, then translates the request into data operations.
PatchOptic is an interface for showing each language-model step only the part of shared workflow state that it needs.
PDEFlow is an autonomous agentic framework that turns user-level ODE and PDE descriptions into solver-backed neural-operator pipelines.
PentestGPT is a security-testing assistant that helps plan and reason through penetration-testing work.
Perplexity is an answer engine that searches the web, writes a response, and shows the sources used to support it.
Comet is Perplexity's browser, pairing web navigation with an assistant that can interpret pages and help complete online work.
Perplexity Pro is a paid AI search service that investigates questions and links its answers to current web sources.
Pipecat is an open-source framework for building real-time voice and multimodal conversational agents.
PLACEMEM is a prototype control system for deciding what a long-lived agent stores, retrieves, updates, and forgets over time.
Plasmate turns a web page into compact structured data so an agent can understand and operate the page with less context.
PlayCode Agent turns plain-English requests into websites that can be built and previewed in a browser.
Playwright MCP is a server that gives AI agents Playwright tools for inspecting and operating web pages.
PolyAI builds enterprise voice agents for natural, multi-turn customer conversations in service industries.
PR-Agent is an open-source tool that describes, reviews, and suggests improvements to pull requests.
Prism Scanner checks agent skills, plugins, and Model Context Protocol servers for risky behavior before and after installation.
Prismata is a defense enforcing contextual least privilege for web agents, constraining both what the agent sees and what it can do.
ProjAgent is a repository-level code generation system that introduces procedural similarity as an explicit retrieval signal.
Prompt Coach is an agentic tutor that helps developers learn how to craft high-quality code-generation prompts through Socratic guidance embedded in-flow within their IDE.
Prompt-to-Paper is a multi-agent framework that directly addresses this evaluation gap through three integrated innovations.
Promptfoo is an open-source testing tool for comparing prompts, checking model and agent behavior, and running security tests.
Pydantic AI is a Python framework that applies Pydantic's validation model to agents, tools, structured outputs, and tests.
Qodo reviews code changes with repository context and checks whether a pull request meets its stated purpose.
Specialist security-operations AI agents for alert triage, incident investigation, threat hunting, vulnerability ranking, and access review across federated data.
Ragas is an evaluation toolkit for testing retrieval, question answering, tool use, and other language-model application behavior.
RAGFlow is an open-source retrieval engine for building grounded question-answering applications and agent workflows.
Rasa is an open-source platform for training and self-hosting conversational assistants with natural-language understanding.
Relevance AI is a no-code platform for building agents that support sales, customer service, and research work.
Replit Agent turns a prompt into a full-stack application and can deploy the finished project from Replit.
ResearchStudio-Reel is a set of connected skills that turns one paper extraction into an editable poster, talk video, and bilingual blog.
Retell AI provides infrastructure for multilingual voice agents that conduct telephone conversations.
Rivet is a visual builder for arranging AI prompts, tools, and decisions into a workflow.
RLVP is a training recipe for real-world agents that penalizes unsafe actions along the way while rewarding a successful final outcome.
RooCode is an open-source coding agent with structured modes for planning, editing, and debugging work.
Salesforce Agentforce is a platform for building and supervising business agents that use Salesforce data, rules, and customer workflows.
Salesforce combines Einstein predictions with Agentforce agents inside its customer-data and automation platform.
Salesmate uses AI to summarize calls and help sales teams qualify leads.
SAP Joule Studio lets organizations create agents and skills that work with SAP business data and processes.
Semantic Kernel is Microsoft's agent-development SDK for C#, Python, and Java.
ServiceNow AI Agents coordinate work across information technology, employee, and customer-service workflows.
SHIELD is a multi-agent teaching system that turns a coding agent's own reasoning into short learning moments for the developer using it.
Signals CLI collects sales signals such as job changes, funding events, and public interest, then returns structured data for an agent workflow.
SkillFab is a platform for turning a missing agent capability into a reviewed, versioned skill that people and other agents can reuse.
SkillReranker is an inference-time reranking framework for adaptive skill selection.
Skyvern is an open-source browser agent that uses visual understanding instead of relying only on coded page selectors.
SMETRIC is a scheduler that spreads first requests across a model-serving cluster, then routes later requests to machines that can reuse the session's cached context.
Smolagents is Hugging Face's small agent library for systems that plan and act by writing code or calling tools.
SocaSim is an LLM-based multi-agent simulation framework to study Putnam's Social Capital Theory from theoretical blueprint to simulated reality.
SpaCellAgent is an autonomous large language model (LLM) multi-agent framework that automates end-to-end spatiotemporal analysis and narrative generation.
StateFuse is a conflict-aware replicated memory contract built on standard OpSet/CRDT merge.
STORM is a Stanford research system that gathers sources and writes long, reference-style articles from a topic.
Strands Agents supplies an open-source SDK for building and controlling production agent harnesses across models and clouds.
SuperAGI is an open-source platform for creating and running agents locally or in the cloud with reusable toolkits for outside systems.
SWE-Agent is a Princeton research agent for resolving real software issues in GitHub repositories.
Synthflow is a no-code platform for small businesses to build voice agents from templates.
TACTIC-KG is an agentic framework for CSKG construction that decomposes the task into modular, specialized LLM agents responsible for extraction, typing, verification, and curation.
TaskWeaver is Microsoft's code-first agent framework for data analysis and tool-using tasks.
TeamHero is an open-source local system for coordinating several coding agents through a dashboard, shared knowledge, and a visible task lifecycle.
Temporal provides durable execution for agent workflows that must survive failures, waits, and long-running jobs.
TOFFEE is a system for synthesizing high-quality data agent trajectories from given data environments via Monte Carlo Tree Search (MCTS) with adaptive model selection and cross-task prefix reuse.
TopoBrick is a training-free framework for zero-shot building IoT (Internet-of-Things) forecasting.
TRACE is a watermark for an agent's action log that can remain detectable after steps are deleted or the written record is rewritten.
Upsonic is a Python framework for autonomous agents, with controls for tools, tasks, teams, memory, and reliability.
Vercel's v0 turns a written interface request into React components styled with Tailwind CSS.
Vapi provides model-independent, low-latency infrastructure for developers building telephone voice agents.
Visual Action Outcome Reasoning Alignment (VAORA) is a novel reward design that directly addresses both issues.
Visual Inspection of Policies uses recorded videos of an agent's behavior to choose which training tasks the agent should try next.
Vocode is an open-source framework for voice agents built with language models.
Voiceflow is a visual, no-code builder for voice and chat assistants.
WCog-VLA is a novel dual-level World-Cognitive VLA framework that successfully bridges semantic world forecasting with generative world evolution to achieve proactive autonomous driving.
WebSwarm is a progressive recursive delegation framework that jointly constructs task decomposition, recursive expansion, and agent collaboration during inference.
Windsurf is an AI code editor whose Cascade agent can plan changes, edit several files, and remember project context.
WRITER Action Agent carries out business tasks across approved tools while keeping its work inside Writer's enterprise AI platform.
YAWNING TITAN is a graph-based simulation environment for cybersecurity research and defensive-agent experiments.
AutoGLM is Z.ai's open phone-agent model and framework for understanding a mobile screen and taking actions inside apps.
Zapier AI turns natural-language requests into workflows that connect supported applications.
Zapier Agents lets people assign goals to AI workers that can use connected apps and run automations on their behalf.
Zendesk AI Agents classify support tickets, detect sentiment, route requests, and answer common questions.
Zia is Zoho CRM's AI assistant for predictions, sentiment analysis, and voice-driven customer-management work.