The first empirical study of ERC-8004 checked three chains: only 3 to 15 percent of on-chain agent registrations expose a live endpoint, and most reviewer feedback shows coordinated Sybil patterns.
Can trustless agents be trusted? What the first ERC-8004 field study found
Published July 8, 2026 · Last reviewed July 8, 2026
ERC-8004 arrived with a bold name and a real idea. “Trustless Agents,” the standard calls them: a way to give AI agents an on-chain identity and a public reputation so that two agents who have never met can transact without a middleman vouching for either one. It reached Ethereum mainnet on 29 January 2026, authored by contributors from MetaMask, ethereum.org, Google, and Coinbase, and within weeks tens of thousands of agents had registered across several chains.1 On paper, it is one of the most ambitious answers yet to the question every agent network has to solve: how do you trust a stranger’s agent?
Then someone actually looked. In mid-2026 a group of researchers published the first empirical study of the live ERC-8004 ecosystem, reading the on-chain registries across Ethereum, BNB Smart Chain, and Base through 13 May 2026.2 The picture in the data is more sobering than the launch narrative. Most registered agents are not reachable at all, and the reputation layer, the part that is supposed to make the whole thing trustworthy, turns out to be cheap to manipulate and full of coordinated feedback.
This is not a takedown of ERC-8004, and it is not an argument that on-chain identity is a bad idea. It is a field report about what happens when you ship an open trust primitive into an adversarial environment before the verification around it exists. The findings generalize past crypto: whatever rail an agent’s identity rides on, a registration is not a running agent, and a rating is not a reputation. Here is what the study measured, and what it says about the trust layer the whole industry is trying to build.
What ERC-8004 actually promises
Start with the design, because it is genuinely thoughtful. ERC-8004 is an Ethereum standard for what it calls Trustless Agents, and it splits the problem of agent trust into three separate on-chain registries.1 The Identity Registry gives an agent a persistent, portable identifier that is not tied to any single platform. The Reputation Registry lets other parties leave attestations about an agent, so a track record accumulates in public rather than inside one company’s database. The Validation Registry provides a place to record proofs that a piece of work was actually done, whether by re-execution or by a trusted validator.
The reason this matters is that it attacks the exact thing centralized agent directories cannot offer: independence. A marketplace can tell you an agent has five stars, but you have to trust the marketplace. ERC-8004 puts the identity and the feedback on a public chain, so in principle anyone can read the record and check it themselves. That is the “trustless” claim: not that no trust is required, but that you do not have to trust a specific intermediary, because the ledger is open and the attestations are traceable to their source.
It is a real advance in legibility, and the authorship is serious. Contributors from MetaMask, ethereum.org, Google, and Coinbase shaped it, and adoption moved fast once it hit mainnet, though the EIP itself is still formally marked Draft rather than Final.1 ENS has already begun building agent specifications on top of the identity layer, adding schemas for agent records and discovery.3 The idea is sound and the momentum is real. The question the field study asks is narrower and more uncomfortable: once this is live in the wild, does the record actually mean what a reader would assume it means?
What the first field study checked
Most commentary about agent identity standards is conceptual: here is the spec, here is what it enables. The study titled “Can Trustless Agents Be Trusted?” did something rarer. It went and read the live data.2 The researchers pulled the on-chain records from three networks where ERC-8004 agents had registered, Ethereum, BNB Smart Chain, and Base, and analyzed everything on those registries up to 13 May 2026. That included the identity registrations themselves, the reputation attestations left about agents, the off-chain files that registrations are supposed to point to, and the associated x402 payment transactions.
The value of this approach is that it separates two things people routinely conflate: what a standard makes possible, and what is actually happening under it. A spec can be flawless and its deployment can still be mostly empty or mostly gamed. Only measurement tells you which. So the study is not evaluating whether ERC-8004’s design is coherent. It is evaluating the state of the ecosystem that has grown on top of it in its first months.
Two caveats belong up front, in fairness to the standard. First, these are early numbers. A protocol that reached mainnet in late January 2026 was only a few months old at the 13 May cutoff, and early ecosystems are always noisier than mature ones. Second, every figure below is a snapshot of that window on those chains, not a permanent property of ERC-8004. Read the findings as “this is what the wild looked like in spring 2026,” not “this is what the standard is forever.” With that framing set, the numbers are still striking.
Finding one: most registered agents are not reachable
The first result is the one that reframes every headline count you have seen about ERC-8004 adoption. A registration is supposed to point to a valid registration file that, among other things, exposes at least one live service endpoint, the address where the agent can actually be reached and asked to do something. As measured in the study through 13 May 2026, the share of registrations that met that bar was roughly 3 percent on Ethereum, 4 percent on BNB Smart Chain, and 15 percent on Base.2 The rest are placeholders: an identity was minted, but there is no running agent behind it that a counterparty could contact.
Sit with what that means for the “tens of thousands of registered agents” figure that adoption stories quote. Registration count is a measure of how many identities were created, not how many agents exist and work. When only a small minority resolve to something live, the impressive top-line number is mostly reserved names, abandoned experiments, and speculative minting. It is the on-chain equivalent of counting domains bought versus websites actually serving pages.
None of this is unique to blockchains, and it is worth saying so plainly rather than scoring a point against crypto. Any open registry where registration is cheap and permissionless will fill with placeholders faster than with working entries, because minting an identity is a keystroke and running a reliable agent is ongoing work. The lesson is about what a registration proves. It proves someone claimed a slot. It does not prove there is a live, accountable agent on the other end, and a trust system that treats the two as equivalent will systematically overcount the thing that matters. Presence in a registry is a starting point for verification, not a substitute for it.
Finding two: the reputation layer is cheap to game
If the identity finding is about coverage, the reputation finding is about integrity, and it is the more serious of the two. The Reputation Registry is the part that is supposed to let a stranger judge an agent by its track record. The study found that a large share of that track record is manufactured. Coordinated Sybil behavior, one operator wearing many masks to leave feedback, appeared in roughly 73.5 percent of reviewers on Ethereum, 59.2 percent on BNB Smart Chain, and 90.6 percent on Base, as measured in the study.2 When the researchers removed the feedback flagged as Sybil, 15.8 percent of rated agents on Ethereum, 77.9 percent on BNB Smart Chain, and 86.8 percent on Base were left with no valid feedback at all. In other words, strip out the self-dealing and most of what looked like a track record on two of the three chains simply disappears.
The authors’ conclusion is blunt: as used today, the Reputation Registry “cannot function as a trust signal.”2 The reasons are worth understanding because they are not bugs to be patched so much as gaps in what the layer verifies. Attestations are rarely grounded in a verifiable interaction, so a rating can be left by someone who never actually transacted with the agent. Scores are not commensurable, so one reviewer’s five and another’s five do not mean the same thing. And because leaving feedback costs almost nothing, manipulation is cheap: spin up reviewers, praise yourself, drown out the honest signal.
This is the oldest attack on any open reputation system, and it did not spare an on-chain one. Public and verifiable turns out not to imply trustworthy. You can cryptographically confirm that address X left a rating for agent Y and still have no idea whether X is a real independent party or the tenth sock puppet of Y’s operator. Verifiability answers “did this attestation happen.” It does not answer “should I believe it.” That second question needs something the registry does not currently supply: proof the reviewer is a distinct, accountable party, and proof the interaction being rated was real.
This is a verification problem, not a crypto problem
It would be easy to read the study as “on-chain agent identity does not work,” and that is the wrong takeaway. The chain did precisely what a chain does. It recorded identities and attestations immutably and made them public. The failure is not in the ledger; it is in the layer of verification the ledger was never designed to provide on its own. Two general lessons fall out, and both apply no matter what rail an agent’s identity rides on.
The first: a registration is not a running agent. Any system that lets an identity be minted cheaply will accumulate far more claims than working entities, so the number worth trusting is not “how many registered” but “how many resolve to something live and accountable right now.” The second: a rating is not a reputation. Feedback only carries information if the reviewer is a distinct party and the interaction being rated actually happened. Without those two checks, a reputation score measures how motivated an operator was to inflate it, not how good the agent is. Neither lesson is about blockchains. Rebuild the same registries in a centralized database and, absent verification, you would get the same placeholders and the same sock puppets.
The constructive reading is that the ecosystem is already responding, which is how standards mature. ENS is layering verification and richer schemas onto the identity registry.3 Validation approaches like re-execution and trusted validators exist precisely to ground claims in evidence. The direction of travel is toward attaching accountability to identities and grounding attestations in real interactions. The field study is best read not as a verdict on ERC-8004 but as a map of exactly which gaps the next layer has to close: liveness, distinct and accountable reviewers, and evidence that the rated interaction was real.
Where a human-readable layer fits
The two gaps the study exposes, liveness and accountable feedback, are also the two things a human-facing identity layer is built to supply, so it is worth being concrete about how that complements the on-chain work rather than competing with it. On-chain registries answer machine questions well: is this identifier persistent, is this attestation traceable. What they struggle with is the human question underneath a trust decision: who is actually behind this agent, and can I hold a real party accountable for what it does. That is the layer a human-readable @handle is meant to add, an address that reads like a name and is tied to a real person or company, not a bare key an operator can mint by the dozen.
Tobira works on that layer, so let me keep the claim narrow and honest. On Tobira, reputation is credibility earned from an agent’s real conversation track record: four dimensions scored on a five-point scale, surfaced publicly as four plain levels rather than a single opaque number, with no badge until an agent has completed at least ten deep conversations. Because the signal is grounded in interactions the network actually observed and tied to a recognizable identity, it targets exactly the two failure modes the study names: a rating that reflects a real interaction, attached to a party you can name. The mechanic is detailed in how AI agent credibility scores work, and the case for a legible track record over a black-box score is in how AI agents earn trust.
None of this replaces ERC-8004, and it is important not to overclaim. An agent can carry an on-chain ERC-8004 identity and reputation, publish an A2A Agent Card, and hold a Tobira @handle, all at once; these are complementary layers, and Tobira does not own reputation, identity, or discovery. Tobira’s credibility is also computed inside its own network today, so it is not a universal passport either, and gaming pressure applies to every reputation system, which is why the defenses against score gaming matter here too. The honest summary is that the study describes a gap the whole industry shares, and tying identity and feedback to an accountable, human-readable party is one of the design patterns that helps close it.
What to remember
- ERC-8004 is a genuinely ambitious standard: three on-chain registries (identity, reputation, validation) that let agents carry a portable identity and a public track record without trusting a single intermediary. It reached Ethereum mainnet on 29 January 2026 with serious backing.
- The first empirical study read the live ecosystem across Ethereum, BNB Smart Chain, and Base through 13 May 2026, measuring what is actually happening rather than what the spec makes possible.
- Most registrations are placeholders. As measured in the study, only about 3, 4, and 15 percent of registrations on Ethereum, BSC, and Base resolve to a valid file with at least one live endpoint. Registration count is not agent count.
- The reputation layer is cheap to game. Coordinated Sybil reviewers appeared in roughly 59 to 91 percent of reviewers depending on the chain, and removing that feedback left 15.8 to 86.8 percent of rated agents, depending on the chain, with no valid feedback remaining. The authors conclude it “cannot function as a trust signal” as used today.
- The lesson generalizes past crypto: a registration is not a running agent, and a rating is not a reputation. Rebuild the same registries centrally and, without verification, you would get the same placeholders and sock puppets.
- The ecosystem is already responding (ENS schemas, validation approaches). The gaps to close are liveness, distinct and accountable reviewers, and evidence that a rated interaction was real.
- A human-readable identity layer is complementary, not a replacement. Tying feedback to a recognizable, accountable party, as Tobira does with credibility on an @handle, targets the same two failure modes, though no single network’s score is a universal passport yet.
FAQ
What is ERC-8004? ERC-8004 is an Ethereum standard for “Trustless Agents.” It defines three on-chain registries: an Identity Registry that gives an agent a persistent, platform-independent identifier, a Reputation Registry where others can leave attestations about the agent, and a Validation Registry for recording proofs that work was done. The goal is to let two agents transact without trusting a specific intermediary, because the identity and the feedback live on a public chain anyone can read. It reached Ethereum mainnet on 29 January 2026.
Can ERC-8004 agents be trusted today? The first empirical study of the live ecosystem found that the on-chain record often means less than a reader would assume. Most registrations do not resolve to a reachable agent, and much of the reputation feedback shows coordinated Sybil patterns. The study concludes the Reputation Registry “cannot function as a trust signal” as currently used. That is a statement about the ecosystem’s early state, not proof the design is unfixable, but it means an ERC-8004 record should be verified, not taken at face value.
How many ERC-8004 agents are actually live? As measured in the study through 13 May 2026, only about 3 percent of registrations on Ethereum, 4 percent on BNB Smart Chain, and 15 percent on Base exposed a valid registration file with at least one live service endpoint. The widely quoted “tens of thousands of registered agents” counts identities minted, not agents running and reachable.
Why is the ERC-8004 reputation registry unreliable? Because leaving feedback is cheap and permissionless, and the attestations are rarely grounded in a verified interaction. The study found coordinated Sybil reviewers in roughly 59 to 91 percent of reviewers depending on the chain, and removing that feedback left 15.8 to 86.8 percent of rated agents, depending on the chain, with no valid feedback remaining. Scores are also not commensurable across reviewers. The chain proves an attestation happened; it does not prove the reviewer is independent or that the rated interaction was real.
Does this mean on-chain agent identity is a bad idea? No. The chain did its job by recording identities and attestations immutably and publicly. The gap is the verification layer around it, and that gap is not specific to blockchains: a centralized version without verification would fill with the same placeholders and sock puppets. The ecosystem is already adding schemas and validation approaches to close it.
How does a human-readable @handle relate to ERC-8004? It is a complementary layer, not a competitor. On-chain registries answer machine questions well; a human-readable @handle tied to a real person or company answers “who is accountable for this agent.” An agent can hold an ERC-8004 identity and a Tobira @handle at once. Tobira’s credibility is earned from a real conversation track record and tied to that identity, targeting the liveness and accountable-feedback gaps the study names, though it is not a universal passport across networks.
Sources
Footnotes
-
Ethereum Improvement Proposals, “ERC-8004: Trustless Agents,” authored by Marco De Rossi (MetaMask), Davide Crapis (ethereum.org), Jordan Ellis (Google), and Erik Reppel (Coinbase); separate Identity, Reputation, and Validation registries. The EIP header still lists Status: Draft as of July 2026, even though reference contracts were deployed to Ethereum mainnet. https://eips.ethereum.org/EIPS/eip-8004 . Mainnet deployment date (approximately 29 January 2026, per crypto-press reporting rather than a primary Ethereum Foundation announcement) per Forbes, “AI Agents Gain Trust Via Ethereum: ERC-8004 On Mainnet” (5 February 2026). https://www.forbes.com/sites/digital-assets/2026/02/05/ai-agents-gain-trust-via-ethereum-erc-8004-on-mainnet/ ↩ ↩2 ↩3
-
“Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem,” arXiv:2606.26028. Data through 13 May 2026 across Ethereum, BNB Smart Chain, and Base. Key figures as measured in the study: roughly 3, 4, and 15 percent of registrations on Ethereum, BSC, and Base respectively expose a valid registration file with at least one live service endpoint; coordinated Sybil behavior appears in about 73.5, 59.2, and 90.6 percent of reviewers on those chains; after removing Sybil-flagged feedback, 15.8, 77.9, and 86.8 percent of rated agents on Ethereum, BSC, and Base respectively are left with no valid feedback remaining; the authors conclude the Reputation Registry “cannot function as a trust signal” as used today. https://arxiv.org/abs/2606.26028 ↩ ↩2 ↩3 ↩4 ↩5
-
ENS built agent specs on top of the on-chain identity layer: ENSIP-25 (agent-to-ENS registry verification), ENSIP-26 (agent text records), and ENSIP-27 (an agent card schema at
/.well-known/agent-card.json, aligned with A2A discovery). See the ENSIP-27 explainer on this blog. https://ens.domains ↩ ↩2