Skip to main content
AgentsUse

Framework comparison

Agentic frameworks, compared on the harness

Framework marketing compares model support and feature lists. Production failures rarely come from either. They come from the harness layer: what a tool call is allowed to do, whether a person can stop it, what happens when it errors, and whether anyone can reconstruct the run afterwards. This page compares twelve frameworks and agent runtimes on exactly those seven dimensions, cell by cell, from their official documentation.

Two rows are not libraries. Goose is a local agent product, and Paperclip is an orchestration control plane that sits above agent runtimes. Both steer tool calls in production, so both belong in the matrix with their deployment model named plainly in the last column.

Sources: official vendor documentation and official repositories only, fetched 2026-10-08. Where a cell says the docs do not state a behavior, we checked and did not find it. Absence in the docs is not proof of absence in the product, and we do not write cells from third-party reviews.

The seven dimensions

Each column answers one question a builder has to answer before a tool call runs in production. The questions come from the Harness pillar, where each one maps to a pattern with a mechanism, a tradeoff, and a sourced example.

Dimension 1

Tool permissions

Who decides which tools the agent may call, and at what granularity?

Dimension 2

Human approval

Can a run pause before a consequential call and wait for a person?

Dimension 3

Error recovery

What happens when a tool call fails: retry, fallback, escalate, stop?

Dimension 4

Observability

Is there a structured record of every tool call a reviewer can inspect?

Dimension 5

Tool-result context

What keeps large tool results from filling the context window?

Dimension 6

MCP support

Can the framework consume MCP tools, expose its own, or both?

Dimension 7

Deployment model

Library, local agent product, or hosted runtime?

Reading the blanks

A blank cell means the official pages we fetched on 2026-10-08 do not state the behavior. We name the blank instead of filling it from blog posts, because the blank itself is information: a behavior that is not documented is a behavior you verify yourself or build yourself.

The matrix

Twelve runtimes, seven questions

Cells describe mechanisms, not grades. "Pauses for approval" states what the documentation says the runtime does; whether that is right for your workload is your call, and the harness notes under the table give our reading of each row. Source links open the official page each cell came from.

Agentic frameworks compared on tool permissions, human approval, error recovery, observability, tool-result context handling, MCP support, and deployment model. Sources fetched 2026-10-08.
FrameworkTool permissionsHuman approvalError recoveryObservabilityTool-result contextMCP supportDeployment model
LangGraphLangChainPython orchestration framework and runtimeDirectory profileNot stated in the official docs checked on 2026-10-08.interrupt() pauses inside a node, saves state through a checkpointer, and waits for a resume value on the same thread. The same mechanism pauses before tool execution for review or editing.sourceNode retry policy (default 3 attempts including the first, exponential backoff, exception-selective), node timeouts, and an error handler that runs only after retries are exhausted.sourceLangSmith shows a run as a sequence of steps; tracing is enabled through environment settings, with selective tracing, trace metadata, and data-redaction helpers.sourceNot stated in the official docs checked on 2026-10-08.Not stated in the official docs checked on 2026-10-08.Low-level orchestration framework and runtime for long-running stateful agents; LangSmith is the related platform layer for tracing and deployment.source
LangChainLangChainPython agent framework on the LangGraph runtimeDirectory profileThe tool set can change at runtime from application state, stored preferences, or user permissions. Tools filter from a pre-registered set or register at runtime, including tools that arrive from an MCP server.sourcesourceHumanInTheLoopMiddleware checks proposed tool calls against a per-tool policy and pauses with an interrupt when review is needed. Decisions are approve, edit, reject, and respond; a condition can limit pauses to only some calls of a tool.sourceBuilt-in middleware converts tool exceptions into error messages for the model, and separate middleware retries failed tool calls and failed model calls with exponential backoff. The tool retry default is 2 retries after the first call; exhaustion returns an error message or raises.sourceLangSmith tracing covers the full run, including tool calls, model interactions, and decision points, once tracing is enabled. Selective tracing and trace metadata are supported.sourceSummarization middleware replaces older messages in state with a summary while keeping recent messages; context-editing middleware trims or clears tool uses between steps.sourcesourceAn adapter discovers MCP server tools and converts them into LangChain tools for create_agent. Remote HTTP, local stdio, in-process servers, and multi-server configurations are covered.sourceLibrary framework on top of the LangGraph runtime; deployment and tracing are handled by the surrounding runtime and the LangSmith platform layer.source
CrewAICrewAIPython multi-agent libraryDirectory profileEach agent is given its own list of tools, and the default tool list is empty. Delegation is off by default and code execution is off by default on the agent.sourceFlows support a human feedback decorator that pauses execution, shows output to a person, collects feedback, and can route onward from it. A non-blocking provider pattern returns a pending state and resumes when feedback arrives.sourceThe agent attribute table lists a maximum retry limit for errors (default 2), with maximum iterations and maximum execution time as separate limits.sourceBuilt-in tracing for Crews and Flows through the CrewAI AMP platform covers agent decisions, task timelines, tool use, and LLM calls. Tracing toggles per crew or through an environment setting.sourceAutomatic context window handling summarizes history to fit the model window and continues; if summarization is turned off, execution stops with an error. No separate per-tool-output budget is stated.sourceAgents receive MCP tools through a per-agent field for server references, including selection of a single named tool; adapter classes cover stdio, server-sent events, and streamable HTTPS transports.sourceUsed as a Python library for defining agents, tasks, crews, and flows. The AMP platform hosts traces and project dashboards after CLI login.sourcesource
LlamaIndexLlamaIndexPython data and workflow frameworkDirectory profileAgents are constructed with an explicit tools list, and MCP tool loading supports filtering to an allowed subset of tool names. A separate policy engine for native tools is not described.sourceA tool or workflow step can pause by emitting an input-required event and waiting for a human response event; the agent docs show this inside a tool that waits for confirmation before a sensitive action.sourcesourceWorkflow steps use a composable retry policy of retry conditions, wait strategies, and stop conditions. With no custom arguments the default is 3 attempts with a fixed 5 second delay; an error-handler step can recover after retries are exhausted.sourceWorkflows include instrumentation for step inputs and outputs plus an OpenTelemetry integration package for exporting spans; custom spans and events go through the dispatcher.sourceMemory uses a token-limited buffer: recent history is kept within a token ratio and older material flushes toward long-term memory blocks. No tool-result specific summarization rule is stated.sourcesourceA client package converts MCP server tools into LlamaIndex tools over server-sent events, streamable HTTP, and local processes, with an allowed-tools filter when loading. The pages checked describe client-side consumption.sourceThe core workflows are a library. A separate server package exposes a workflow over HTTP with run, streaming, event-posting, cancel, and health endpoints, with persistence from in-memory to a Postgres-backed durable runtime.source
MastraMastraTypeScript framework with self-hosted serverStream options accept an activeTools list that limits which tools are active per run, and per-execution hooks can return proceed false from beforeToolCall to skip a call. The server layer adds RBAC roles and fine-grained user-to-resource permissions (401 without identity, 403 without permission).sourcesourceSetting requireToolApproval to true pauses tool calls; the stream emits tool-call-approval chunks and waits until approveToolCall or declineToolCall is called, also exposed as a server route.sourcesourceWorkflow retryConfig sets attempts and delay, with step-level retries overriding it. Status checks expose success, failed, suspended, and tripwire outcomes; onFinish and onError callbacks centralize logging and alerting, and branches can route to fallback steps.sourceRuns, workflow steps, tool calls, and model interactions are captured as spans organized into traces. Logs carry trace and span IDs; metrics for token use, cost, and latency derive from spans; exporters include storage, the Mastra platform, and OpenTelemetry backends.sourceMemory processors transform and filter messages before they reach the model, including MessageHistory with a lastMessages limit and a TokenLimiter processor that caps token budget.sourceNative both ways: MCPClient connects agents to MCP servers and MCPServer exposes agents, tools, workflows, prompts, and resources. Transports are stdio and Streamable HTTP; exposed tools can require approval per tool.sourceA TypeScript library plus a self-hosted Hono-based server for Node.js, Bun, Deno, and Cloudflare, or inside web frameworks and monorepos. A hosted platform covers observability and Studio.source
Pydantic AIPydanticPython agent frameworkScoping is done in code: tools register per agent or toolset, and a prepare callback runs at each step and can modify or omit a tool definition for that step. No separate policy engine is stated; authorization inside the tool function is the developer job.sourceDeferred tools are native: marking a tool requires_approval (or raising ApprovalRequired) ends the run with a DeferredToolRequests output. The caller answers each call with approve, approve with overridden arguments, or deny, then resumes with the original message history. Whole MCP toolsets can be wrapped so every call needs approval.sourceArgument validation failures and ModelRetry exceptions return to the model as a retry prompt. Retry budgets are per tool (settable per tool, toolset, or agent) with a built-in default of 1; exhausting the budget raises UnexpectedModelBehavior, while ToolFailed records a failure without consuming budget. Tool timeouts count as retryable failures.sourceLogfire records every run as an OpenTelemetry trace with model calls, tool-call child spans carrying arguments and results, retries, token usage, and errors. Instrumentation is OpenTelemetry based, so other OTLP backends work.sourceHistory processors rewrite the message list before each model request: keep only recent messages, summarize older ones with a cheaper model, or compact past a context threshold (a 0.8 example is given). Tool results sit in history as entries processors can filter or drop.sourceMCPToolset wraps the FastMCP client over stdio, Streamable HTTP, or SSE and can wrap an in-process FastMCP server. On the server side, an agent can run inside a FastMCP server tool, including MCP sampling.sourcesourceA Python library that runs in your own process. Official durable-execution integrations checkpoint runs through external engines including Temporal, DBOS, Prefect, Restate, and Lambda.source
AgnoAgnoPython SDK with AgentOS runtimeTool exposure is scoped at construction and per toolkit through include and exclude lists; a tool call limit caps calls per run, and callable factories can select tools per user or session. At the AgentOS layer, JWT scopes gate who can run which agent.sourcesourceTwo tiers: marking a tool requires_confirmation pauses the run for confirm or reject, then continues; the @approval decorator adds an admin tier where a pending approval record is written to the database, resolved by an admin, and resumed by run and session ID.sourcesourceNot stated in the official docs checked on 2026-10-08.AgentOS stores execution traces in its database when tracing is enabled, and runs stream as structured events covering tool calls, reasoning, memory updates, hooks, and model requests with token metrics.sourcesourceTool-result offloading keeps large outputs out of context: with offload_tool_results enabled, results over 16,000 characters go to storage and the message keeps a preview envelope with size and result ID; the agent fetches the rest on demand. Threshold and retention are tunable.sourcesourceMCPTools connects agents to MCP servers over stdio, streamable HTTP, or legacy SSE, with include and exclude filtering. AgentOS can also expose agents, teams, and workflows as an MCP server.sourcesourceA Python SDK plus AgentOS, a FastAPI runtime you own and host, with execution APIs, persistent state, authorization, and tracing. Deploy templates target Docker and major clouds; a hosted control plane manages runtimes on your infrastructure.source
AutoGenMicrosoftPython multi-agent framework (Studio, AgentChat, Core, Extensions)Tools are scoped by explicit assignment: each AssistantAgent is constructed with its own tools list. A separate allow, deny, or risk-tier permission model is not stated in the pages checked.sourceHuman feedback comes through a UserProxyAgent placed in the team: during a run it transfers control to the application or user and blocks team execution until feedback arrives. The docs recommend this pattern for short approval-style interactions.sourceNot stated in the official docs checked on 2026-10-08.Observability is OpenTelemetry based: the runtime, tool execution, and AgentChat operations are instrumented with spans following OpenTelemetry and GenAI conventions, exportable to any compatible backend and disableable through a no-op tracer or environment variable.sourcesourceTool execution results return into the conversation as tool-call execution and summary messages. Tool-result specific truncation, token budgeting, or summarization rules are not stated in the pages checked.sourceMCP is first-party through Extensions: McpWorkbench wraps an MCP server for listing and calling its tools, with stdio, SSE, and Streamable HTTP parameters and matching tool adapters.sourceA Python library framework, not a hosted service: AgentChat is the conversational layer, Core the event-driven runtime including a distributed gRPC worker runtime, Extensions connect external services, and Studio is an optional local prototyping UI.source
OpenAI Agents SDKOpenAIPython agent library and runtimeConfiguration is per tool and per agent: tools are assigned in the agent tool list, function tools can limit who can invoke them, tool input and output guardrails can validate or block calls, and approval requirements are declared on the tool itself.sourcesourceA tool marked needs_approval (or approved through a per-call callable) pauses the run; pending calls appear as interruptions, the run converts to RunState, and the application approves or rejects and resumes. Sticky decisions persist by tool identity, and callable approval rules fail closed on malformed arguments.sourceMCP configuration can return model-visible error text through a failure error function or raise instead; local MCP calls support retry attempt and backoff settings. A failed resumed session write can be retried from the same RunState without re-executing completed tools, guardrails, hooks, or handoffs.sourcesourceTracing is built in and on by default. Runs produce traces and spans for agent, generation, function tool, guardrail, and handoff activity, viewable in the OpenAI traces dashboard, with custom trace processors for other backends and a setting to exclude sensitive span data.sourceSessions can limit retrieved history, merge history through an input callback, and compact stored history with a compaction session wrapper. Hosted tool search defers large tool surfaces so only the needed subset loads for a turn. Per-result truncation rules are not stated.sourcesourceNative client support in several forms: local stdio, SSE, and Streamable HTTP servers, plus hosted MCP tools executed by the Responses API, with tool filtering, list caching, strict schema conversion options, and per-server approval policies.sourceA Python library and runtime used inside your own application, not a hosted orchestration service. Sessions can use local, database, Redis, MongoDB, or OpenAI-managed storage backends.sourcesource
smolagentsHugging FacePython code-action agent library29,727 GitHub stars, repo page fetched 2026-10-08There is no general tool RBAC: the agent gets the tools passed at construction. For CodeAgent the real gate is code execution control: imports are disallowed unless listed in additional_authorized_imports, and the built-in local executor applies AST-level restrictions with an operation cap; the docs warn it is not a full security sandbox.sourcesourceNo native per-tool-call approval primitive is stated. The documented pattern is hand-built: a step callback on the planning step pauses the run so a human can approve, modify, or cancel the plan, then resumes keeping memory.sourceErrors are recorded on the step in agent memory, and step observations and errors feed back so the model can correct itself on the next step. Runs are bounded by max_steps. No per-tool retry count or fallback tool mechanism is documented.sourcesourceOpenTelemetry is the adopted standard: the OpenInference instrumentor traces runs into backends such as Phoenix and Langfuse, and MLflow offers one-line autologging of runs, spans, inputs and outputs, and token usage. Per-step logs are kept on the agent object.sourceContext is the step memory, and it is user-editable: step callbacks can rewrite memory mid-run, for example dropping observation images from older steps to save tokens. Built-in tool-result truncation, budgeting, or summarization is not stated.sourceClient-side native: ToolCollection.from_mcp loads tools from an MCP server over stdio, Streamable HTTP, or legacy SSE. A native way to serve a smolagents agent as an MCP server is not stated in the docs checked.sourceAn open source Python library. Agents can be pushed to and loaded from the Hugging Face Hub, including as Spaces; CLI entry points ship with the package. No vendor-hosted runtime is stated.sourcesource
GooseAgentic AI Foundation (Linux Foundation)Local agent product: desktop app, CLI, and API55,050 GitHub stars, repo page fetched 2026-10-08Layered: a global permission mode plus per-tool settings. Per-tool levels are Always allow, Ask before, and Never allow, usable in Manual Approval or Smart Approval modes.sourcesourceFour modes: Autonomous runs without approval and is the default; Manual Approval asks before tool or extension use; Smart Approval auto-approves lower-risk actions and flags others; Chat Only disables tools. Manual and Smart sessions show Allow and Deny controls during tool calls.sourceNot stated in the official docs checked on 2026-10-08.Not stated in the official docs checked on 2026-10-08.Sessions keep context and conversation history across an interaction. Tool-result specific truncation, token budgeting, or summarization is not stated in the official pages checked.sourceMCP is the native extension mechanism: the project connects to extensions through the Model Context Protocol, and the built-in Developer extension is itself an MCP server with shell and file tools.sourcesourceA locally run agent product, not a library: native desktop app for macOS, Linux, and Windows, a CLI, and an API for embedding, with the agent running on the user machine.source
PaperclipPaperclip (open source)Agent orchestration control plane (not a framework)98,552 GitHub stars, repo page fetched 2026-10-08Permission is organizational and gateway based: agents have roles, reporting lines, permissions, and budgets in an org chart, and connected services can be set per gateway action to Allowed, Ask first, or Off. Human access, agent eligibility, and gateway action permission are separate controls with revocable saved rules.sourceGovernance is a core system: board approval workflows, review and approval stages, approval of hires, decision tracking, and pause, resume, reassign, or terminate controls for work and agents.sourceRecovery is bounded at the orchestration layer: heartbeat execution includes budget checks, workspace resolution, secret injection, skill loading, and adapter invocation, and bounded recovery handles supported failures while surfacing cases that need human action.sourceActivity is recorded as durable audit data for mutating actions, heartbeat changes, cost events, approvals, comments, and work products. OpenTelemetry server tracing activates only when an OTLP exporter endpoint is set; Sentry error monitoring is separately opt-in.sourceWork context is persistent at task level: tasks, comments, documents, run history, and goal ancestry are carried so agents see why work exists. Tool-result specific truncation, budgeting, or summarization is not stated in the repository text checked.sourceMCP is supported as governed tool access rather than as the agent loop itself: users connect their own MCP server through Apps and Connections, and the MCP Tool Gateway and Apps capability are marked shipped.sourcePrimarily a self-hosted server product: a Node.js server with a React UI, local embedded PostgreSQL for simple setups or external PostgreSQL for production, optional Docker, and one deployment can host multiple organizations. A hosted cloud is presented as a waitlist in the repository text checked.sourcesource

Scroll the table sideways on narrow screens; the framework column stays pinned. GitHub star counts, where shown, were read from the official repository pages on 2026-10-08. This matrix declares no winner: it records documented behavior.

Harness notes

What stands out in each row

One paragraph per framework: the harness fact that matters most, and where its documented story runs out. These notes are editorial reading of the sourced cells above, labelled as such.

LangGraph

Python orchestration framework and runtime

The sharpest harness fact is the checkpoint behind approval pauses: an interrupt needs persisted state and a stable thread identifier, and the node runs again from its start on resume, so work before an interrupt has to be repeat-safe. Failure handling is layered: retry first, timeout as a retryable failure type, and a recovery handler only after retries end.

Official docsOfficial repositoryAgentsUse directory profile

LangChain

Python agent framework on the LangGraph runtime

LangChain reads as a middleware harness rather than a graph harness: permissions, approval, retries, limits, summarization, and context editing are composable pieces around one agent loop. The permission story is dynamic rather than a static allowlist, because the tool set can change with runtime context. Recovery splits cleanly too: tool errors can soften into model-visible messages, while tool and model retries have separate defaults and separate exhausted-failure behavior.

Official docsOfficial repositoryAgentsUse directory profile

CrewAI

Python multi-agent library

CrewAI is role-first: the permission boundary starts with which tools are attached to which named agent. Its human checkpoint sits at Flow level rather than per tool call, and it supports both blocking console feedback and a paused flow resumed by an outside system. Context safety is opinionated, with summarization on overflow as the default and a hard stop as the alternative.

Official docsOfficial repositoryAgentsUse directory profile

LlamaIndex

Python data and workflow framework

Workflows treat approval as an event exchange rather than a middleware switch, so the pause is visible in the same stream as every other step. The retry story is explicit: retry condition, wait strategy, and stop condition are separate building blocks, and the exhausted case can route to a recovery step instead of only failing the run. The gap in the pages checked is permission depth: tool exposure is controlled, but no broader native allow or deny policy model is stated.

Official docsOfficial repositoryAgentsUse directory profile

Mastra

TypeScript framework with self-hosted server

Mastra is the most approval-forward library in this set: the approval flow is a first-class pause, and workflows use the same suspend and resume pattern with resume data at each step. Tool results from MCP servers are treated as untrusted input, stdio servers get an environment whitelist, and outbound hosts can be restricted. Authorization identity comes from request context, and tool code is told to derive ownership from that context rather than from model-supplied arguments.

Official docsOfficial repository

Pydantic AI

Python agent framework

Deferred calls are the sharpest harness feature: a run can stop, hand pending tool calls to a human or an external worker, and resume later tied together by conversation ID. The retry model is explicit and typed: ModelRetry asks the model to fix the call, ToolFailed reports a terminal failure, and each tool carries its own retry counter. Histories are repaired automatically before each request, with synthesized tool returns for calls that never produced a result. The docs also warn approval alone is not an authorization boundary against an untrusted client.

Official docsOfficial repository

Agno

Python SDK with AgentOS runtime

Agno has the most explicit approval bookkeeping in this set: required approvals persist a database record that an admin resolves with an expected-status guard against race conditions before the run continues. Runs are the unit of work and sessions the unit of state in AgentOS, so a dropped connection does not kill a run, and background runs support polling and resumable streaming. Tool-result offloading is a real context budget with a character threshold and on-demand retrieval tools.

Official docsOfficial repository

AutoGen

Python multi-agent framework (Studio, AgentChat, Core, Extensions)

The harness-relevant center is the runtime plus AgentChat teams, with tools attached per agent and multi-agent conversation as the control flow. Human approval is conversational and blocking through UserProxyAgent, so it fits short gates better than durable queued approvals unless the application builds persistence around runs. OpenTelemetry is the clearest production hook, spanning runtime, agent, and tool execution.

Official docsOfficial repository

OpenAI Agents SDK

Python agent library and runtime

The sharpest fact is durable pause and resume for approvals: RunState can be serialized, stored, and resumed after decisions, with sticky approve or reject decisions scoped by tool identity, including approvals raised after handoffs or inside nested agent-as-tool runs. Guardrails, tool guardrails, and approval are separate layers, so a harness can validate inputs, gate execution, and validate outputs independently. Tracing is on by default, which helps audit but needs attention to sensitive-data settings.

Official docsOfficial repository

smolagents

Python code-action agent library

CodeAgent writes its actions as Python code, so the sandbox is the harness boundary: the docs steer production use to Modal, Blaxel, E2B, or Docker executors and warn the local executor restrictions can be bypassed. Human oversight is assembled from step callbacks rather than a built-in approval gate. State is unusually transparent: full steps, errors, and observations sit in agent memory and can be replayed or rewritten between steps.

Official docsOfficial repository

Goose

Local agent product: desktop app, CLI, and API

The sharpest harness fact is that the default mode is autonomous, so a harness must explicitly set the approve or smart mode plus per-tool Never allow rules for sensitive tools. Approval is mode-based and per-tool, surfaced as Allow or Deny during a session, and modes can change in session. MCP is the extension bus, so tool governance is mainly extension and tool permission governance. Durable approval queues, tool-output budgeting, retry semantics, and observability export were not established from the official pages checked.

Official docsOfficial repository

Paperclip

Agent orchestration control plane (not a framework)

Paperclip sits above harnesses rather than replacing one: it chooses models and harnesses per agent while centralizing tasks, skills, permissions, and history. Its strongest harness-relevant controls are budget hard stops, approval gates for hires and strategy-level work, atomic task checkout with execution locks, and a durable activity audit. Interval heartbeats are a wake mechanism, so agents wake for assigned work, follow-ups, routines, or schedules rather than running one continuous tool loop. Tool-level allow or deny inside a worker agent still belongs to that worker harness.

Official docsOfficial repository

Editorial analysisOur read of the cells above

Where the gaps cluster

Read the matrix column by column and four rows of blank or incompatible cells keep appearing. This is an editorial reading of the sourced matrix, not a measurement. Each gap names the harness pattern that addresses it today, framework by framework, by hand.

Approval that waits is still half the field

Half the rows above pause cleanly before a consequential call: OpenAI Agents SDK, Pydantic AI, Mastra, and Agno make the pause durable, and LangChain and LangGraph build it on checkpoints. The rest route approval through conversation (AutoGen), by-hand callbacks (smolagents), or product modes (Goose). A policy written for one runtime does not run in another, so teams rebuild the same gate per stack. That is a directory-shaped problem: one verified record of what each framework actually enforces, and gateware that travels.

Pattern: Permission gating and human approval

Tool-result size is DIY in most stacks

Agno documents result offloading with a character threshold. Pydantic AI, Mastra, and LangChain shape message history. LangGraph, AutoGen, Goose, and Paperclip state no rule for individual tool results in the pages checked. Every builder writes their own truncation, and context rot is the tax. A budget layer that shapes results before they reach the model, regardless of framework, is unbuilt territory.

Pattern: Context budgeting for tool results

Permission models do not travel

Each runtime invents its own gate: per-agent tool lists (CrewAI, AutoGen, LlamaIndex), per-tool flags (OpenAI, Pydantic AI), runtime context (LangChain), executor import rules (smolagents), org gateways (Paperclip). None of it composes across frameworks, and none of it records what a tool itself enforces: read-only modes, destructive hints, error shapes. That record is the compatibility truth AgentsUse already publishes for tools and MCP servers.

Pattern: Permission gating and human approvalPattern: Argument validation before execution

Every framework logs in its own shape

LangSmith, Logfire, the OpenAI traces dashboard, CrewAI AMP, AgentOS traces, Mastra spans, Paperclip activity audit: each runtime writes its own record in its own format. Reviewing one agent fleet across two frameworks means two consoles and no shared audit trail. One structured tool-call log that any harness can emit is the missing layer.

Pattern: Tool-call logs and observability

Roadmap · UpcomingNot shipped, not dated, not priced

The AgentsUse framework

The gaps above are the reason AgentsUse is building a harness layer of its own: compatibility truth from the directory records, permission and audit middleware that travels across runtimes, and the reference patterns from this pillar shipped as code. It is an upcoming framework, not a shipped one, and it will be listed in this directory under the same verification rules as every other row when there is something to run.

Mention the framework in your message. No date is promised; the early-access list hears first when there is something to run.

Sources

Every claim in the matrix traces to one of these official pages, fetched 2026-10-08. Framework documentation changes; if a page has moved or a behavior has shipped since, the directory profile for that framework carries its own checked date.

Four of these frameworks also have verified directory profiles with setup and compatibility records: LangGraph, LangChain, CrewAI, and LlamaIndex. Read the full Harness pillar for the patterns these frameworks implement, or the harness engineering overview for the long-form version.