AI agency is the shift from rule-based automation to systems that perceive, plan, and act with tools. Learn the spectrum, how Web3 changes it, real use cases, risks, and how to build or work with agents.

Automation follows rules you write. AI agency pursues goals you set. An AI agent perceives its environment, makes a plan, calls tools or smart contracts, and adjusts based on results, all within limits you define.
This matters for anyone building or operating on Web3, because blockchains make agent actions auditable and enforceable in code, but they do not fix model errors or key management failures. This guide maps the spectrum from automation to autonomy, shows where Web3 agents actually work today, and where human oversight still belongs.
AI-driven agency is the ability of software to act toward a goal with limited step-by-step direction.
NIST frames the current work as agents capable of autonomous actions on behalf of users, with standards needed for interoperability and identity. IBM describes the same core loop: perception, then processing, then action. Perception is what separates an agent from a rule-based program.
This is a spectrum, not a binary. Most systems in production sit in the middle: they automate part of a workflow and ask a person to approve actions that move value or change state.
If you only use simple bots or scripts today, this guide helps you see when an agent is warranted and when a deterministic script is safer and cheaper.
Follows if-then-else logic you wrote. Robotic process automation (RPA) is the classic example: software that replays clicks and API calls to handle high-volume, repetitive work. RPA platforms such as Blue Prism, UiPath, and Automation Anywhere excel at structured data and stable interfaces. They fail when the UI changes or an exception appears, because there is no reasoning step.
Add branching and error handling, but every path still needs explicit code. Useful for well-bounded workflows where you can enumerate cases in advance.
Add prediction without action. A spam filter predicts whether a message is spam. It does not decide what to do next or take action in another system. It outputs a score. A human or a separate program must act on it.
Combine a foundation model with memory, planning, and tool use. A typical pattern documented by IBM and in LangGraph tutorials is a plan-and-execute loop:
Memory matters here. Teams usually build three layers: short-term context for the current task, episodic memory per user or project, and semantic memory grounded by retrieval-augmented generation over a vector store. Good agents summarize and ground facts before they act.
Operate for extended periods without human review. Researchers at the Knight First Amendment Institute describe five levels that help make this explicit: user as operator, collaborator, consultant, approver, and observer. At the lowest level, the agent only acts when invoked. At the highest, the user observes while the agent plans and acts and reports back.
Practitioners working with the NIST AI Risk Management Framework also use a four-tier taxonomy, from fully supervised assistance to full autonomy, to apply proportionate controls. NIST published its AI Agent Standards Initiative in early 2026 and, with the National Science Foundation, is developing open protocols and identity research for human-agent and agent-to-agent interaction. A Federal Register request for information on security considerations for AI agents closed in March 2026, with guidance to follow.
Autonomy is a design choice, not an automatic upgrade. Higher autonomy does not mean a better agent. It means you accept more consequence and need stronger guardrails.
An AI agent usually has most of these, while traditional ML has only one or two:
Operates without per-step prompts when you permit it, but stays bounded by rate limits, spend caps, approval gates, and audit logs.
If a system lacks planning and tool use, it is not an agent. If it cannot adapt after an error, it is a script with a model attached.
Blockchains add three properties that matter for agents and one standard that is now taking shape.Transparent execution.
Every transaction an agent sends is recorded on chain and can be audited. You can see what it called, when, and with what gas price. This helps with debugging and dispute review. Smart contract constraints.
Agents operate inside contracts that enforce rules even if the agent's off-chain logic is flawed. A contract can cap daily spend, restrict which pools can be touched, or require a time lock. This does not prevent key compromise or oracle errors, but it bounds what a compromised agent can do. Composability.
On Ethereum, smart contracts are public and can call each other. The ethereum.org docs describe composability as modularity, autonomy, and discoverability: each contract does one job, runs on its own, and is openly addressable so developers can reuse it. An agent can therefore combine a swap, a lending deposit, and a governance vote in one flow without asking each team for permission. DAOs as coordinators.
DAOs encode voting and treasury rules in smart contracts and enforce outcomes through token-weighted votes. Major DeFi protocols governed this way include Aave, Uniswap, Balancer, and Lido, with Layer 2 scaling via Arbitrum and similar networks. Proposals span technical and economic parameters that are hard for many token holders to evaluate, which creates pressure to use agent assistance for analysis and execution while keeping humans as approvers. ERC-8004: Trustless Agents.
In August 2025, contributors from MetaMask (Marco De Rossi), the Ethereum Foundation (Davide Crapis), Google (Jordan Ellis), and Coinbase (Erik Reppel) proposed ERC-8004. It defines three lightweight per-chain registries:
The proposal requires EIP-155, EIP-712, EIP-721, and EIP-1271. It is minimal by design: it handles identity, reputation, and validation, not payments or messaging, which stay with protocols like A2A, MCP, and x402. As of October 2025 it was in Draft, with prototype work shown ahead of DevConnect in November 2025. The Identity and Reputation registries were deployed to Ethereum mainnet on January 29, 2026, with indexing across multiple EVM chains following. This gives agents a portable identifier and a public history that any client can check before delegating funds or tasks.
Most live systems run at the approver level: the agent drafts and simulates, a person or a contract policy approves value-moving steps.
MEV search and block building. Maximal extractable value (MEV) is the value that can be extracted by including, excluding, or reordering transactions in a block beyond the block reward and base fees, per ethereum.org. The current supply chain involves searchers who scan the public mempool, builders who pack bundles into the most profitable block, and validators who propose it. Bots that do this are early autonomous agents: they ingest low-latency mempool data, parse transactions, detect opportunities such as arbitrage or liquidations, and submit bundles. A 2025 overview estimated annual MEV extraction above $3 billion, with methods including sandwich attacks and liquidation racing, though estimates vary by methodology and window. Mitigations now in use include proposer-builder separation and MEV-aware relays and batch auctions. Chainstack and similar providers note that this game is measured in milliseconds, so infrastructure placement matters.
Yield optimization and liquidity management. Yearn Finance popularized automated yield aggregation by routing deposits to the best available strategy and adjusting as rates change. In early 2026, PancakeSwap released agent tooling to compare opportunities across eight chains, and Uniswap Labs released open-source tooling for agents to handle swaps and liquidity on Uniswap v4, which uses a singleton architecture that reduces contract sprawl. BrahmaFi Morpho agents had locked more than $20 million on Base by mid-2025 on one deployment, a concrete example of capital flowing into AI-managed strategies. Chainstack also publishes a local-first tutorial that runs a full trading loop on Base and Uniswap v4 using Foundry to fork mainnet, plus Ollama for local inference, explicitly marked not for production.
Risk and collateral monitoring. Protocols use agents and simulation services to tune parameters. Gauntlet and Chaos Labs run stress tests for Aave and Compound, updating collateral factors and rate curves more often than governance votes alone would allow. For individual positions, agents watch loan-to-value ratios and add collateral or reduce exposure before liquidation, rather than reacting after the fact.
Governance and RWA workflows. Agent prototypes now summarize proposals, simulate effects, and cast delegated votes through narrow permissions, while real-world asset platforms use agents for valuation checks, eligibility, and payout scheduling. These remain assistive, not replacements for accountable owners.
What is still rare is the observer level: an agent that holds assets, votes, and rebalances for days with no human checkpoint. Technical and policy work to make that safe is in progress, but not standard in production custody.
Helps with:
Does not solve:
| Use case | What an agent does | Practical benefit | Constraint or risk |
|---|---|---|---|
| Algorithmic trading and arbitrage | Scans DEXs and mempool feeds, batches swaps, simulates outcomes before sending | Captures short-lived spreads at any hour, reduces missed opportunities | Slippage, gas spikes, and MEV competition can erase edge; needs fast RPCs |
| Liquidity provision on a DEX | Sizes positions, rebalances ranges on concentrated pools, harvests fees | Keeps exposure in target range without manual clicks | Impermanent loss remains, and poor range choices lock capital in low-fee zones |
| Yield farming | Tracks APRs, reward schedules, and protocol risk, moves capital when risk-adjusted return improves | Saves attention across many pools and batches harvests to cut gas | Protocol risk, depeg, and incentives that end abruptly; batch savings are real but limited |
| Bridge operation | Monitors liquidity and finality across chains, submits cross-chain messages | Automates rebalancing that is tedious for a person | Bridge exploit history is significant; limits and circuit breakers are needed |
| Validator and infra operation | Keeps node up, manages attestations, handles upgrades | Improves liveness and reward capture | Slashing for misconfiguration; staking capital is locked |
| Portfolio management for users | Rebalances by target weights, tops up collateral before liquidation | Prevents forced liquidations in volatile periods | Still needs clear risk limits and a pause button you can reach quickly |
| DAO operations | Drafts treasury moves, summarizes proposals, executes approved actions via contracts | Reduces cognitive load for token holders | Vote delegation and proposal complexity remain; agent advice needs provenance |
Alignment.
An agent that optimizes for short-term return can take actions you would reject, such as adding borrowed capital in a thin market. This reward hacking problem is well documented for LLM-based agents. Treat every objective and metric as brittle until proven otherwise. Transparency and explainability.
Deep learning models often operate as black boxes. You can see what transaction an agent sent, but not always why it chose that pool or price. Logging prompts, tool calls, retrieved context, and model version with each action helps, but does not make every decision fully explainable. Out-of-distribution failures.
When an agent meets a situation it did not see in testing, such as a new pool type or an oracle delay, it may fail in unsafe ways. Safe defaults are to do nothing, require approval, or shrink position, not to improvise with funds. Security.
Agents that control assets are high-value targets. Threats documented in 2025 include prompt injection, tool misuse, permission overreach, and agent hijacking where third-party content steers the agent. NIST's 2025 technical blog on strengthening agent hijacking evaluations and the OWASP Agentic Top 10 both emphasize testing with adversarial tool outputs. Keep keys scoped, use ephemeral approvals, and store signing in hardware or a policy engine, not in plain agent memory. Accountability.
If an agent causes harm, legal responsibility sits with the deployer and operator who granted authority, not with the model. Map every permission to an owner, log the delegation chain, and define who can pause, revoke, or roll back an action. Standards and frameworks that teams cite here include the NIST AI Risk Management Framework functions Govern, Map, Measure, and Manage, and the EU AI Act requirements for logging, data governance, documentation, and human oversight for covered high-risk systems. Scalability and systemic risk.
If many agents chase the same signal, exits crowd at the same time. Network congestion blocks clean exits, slippage rises, and correlated liquidations follow. This herd effect already appears in manual farming; faster agents can make it larger rather than smooth it. Trust. Building trust requires both auditable execution and verifiable constraints. A public log alone does not make an agent trustworthy if no one can constrain its spend or verify its data sources.
Always-on agents push markets toward faster price adjustment, which can reduce small arbitrages while raising intraday volatility during stress. The net efficiency effect depends on diversity: many uncorrelated strategies dampen shocks, many correlated ones increase them.
Projections vary widely, which is itself a signal to stay conservative. G2's 2025 AI Agent report reported that 57 percent of surveyed companies had AI agents in production and 78 percent planned to increase agent autonomy. MarketsandMarkets estimated the AI agent market at about $7.84 billion in 2025, rising to over $52 billion by 2030. Treat these as directional survey and forecast data, not as guarantees.
Where an agent helps most:
Where a simpler fix wins:
Start with research to brief to execution for a single protocol, not with a plan to automate everything. 2. **Define tools and least privilege.
List every API and contract the agent may call. Restrict to an allow list, cap spend per day, and require explicit approval for any transaction that moves funds or changes permissions. 3. **Design memory and grounding early.
Decide what is kept in short-term context, what is stored as episodic history, and what is retrieved from a knowledge base via RAG. Ground claims in retrieved documents before acting. 4. **Set the autonomy dial.
Use human-in-the-loop for anything irreversible: treasury moves, bridge transfers, and governance votes. Reserve unattended runs for read-only analysis or sandboxed simulations. 5. **Instrument everything.
Log prompts, tool calls, arguments, costs, latency, and outcomes. Save the transcript so a reviewer can replay why an action was taken. 6. Pilot with shadow review. Let the agent propose for a week while you compare its proposals to what you would have done. Expand scope only after measured accuracy and cost are stable.
Works on planning reliability, provenance, and evaluation of multi-agent behavior.
These roles overlap. In small teams one person covers several, which raises the need for explicit ownership over who can grant and revoke agent authority.
RPA replays fixed steps on a stable interface. A classic trading bot follows a threshold rule you coded. An agent interprets a goal, plans multi-step work, calls tools, and revises its plan based on observations. RPA stays deterministic. An agent adds language reasoning and flexible tool use, at the cost of new failure modes.
Not always. If your workflow fits in context and uses live APIs, you can start without one. Add a vector store and RAG when the task needs multi-session knowledge or grounding across many documents. For chain choice, Base and other low-fee EVM Layer 2 networks are common for prototyping because they keep simulation and experiment costs down.
It gives an agent a portable on-chain identity via an ERC-721 token, plus standard places to record client feedback and independent validation scores. It does not run the agent, move money, or send messages itself. Think of it as a trust directory that other protocols and clients can read before they interact.
Parts of DAO work are assisted by agents, such as summarizing proposals or drafting a treasury sequence, but most DAOs keep voting and high-value execution behind human or contract gates. Research on DAO-AI patterns shows agents can ingest proposal metadata, forum discussion, and voting dynamics to recommend actions, yet accountable owners still set policy and can override.
At minimum: the user goal, the plan, every tool call and argument, retrieved context with source, model and prompt version, gas and cost, the signed transaction hash, and the outcome. On chain, keep only hashes or commitments if privacy matters, and store detail off chain with a pointer. Require that logs be immutable for a retention window.
When its confidence is low, when data sources disagree, when gas is unusually high, or when it encounters a contract or pool it has not seen before. A safe policy is to pause and ask for approval in any of those cases, and to have a global pause that any operator can trigger.What is the most common mistake teams make? Granting broad key permissions before they have defined how to measure task success and before they have tested failure paths. Start narrow, measure accuracy and cost per task, introduce adversarial tests, and only then widen the permission set.
Explore more guides and career playbooks