Agent Readiness Score: Which Layer Are You Testing?

28 August 2026 · updated 06 September 2026 · 2,652 words

Professional header image for industry analysis: Agent Readiness Score: Which Layer Are You Actually Testing?

Most teams deploying AI agents believe they are ready for production long before they actually are. They run a few end-to-end tests, check that the model returns coherent responses, and ship. Then the incidents start rolling in, and suddenly "readiness" becomes a much more complicated concept.

This is where the agent readiness score framework becomes essential. Rather than treating readiness as a binary pass/fail condition, it forces you to evaluate your agent across distinct architectural layers: the model itself, the tool integrations, the orchestration logic, the memory systems, and the external dependencies that tie everything together. Each layer fails in different ways, and each requires a different testing strategy.

In this analysis, we will break down what an agent readiness score actually measures, why most evaluations miss critical failure surfaces, and how to map your current testing coverage against the layers that matter most. By the end, you will have a clearer picture of where your testing is solid, where it is dangerously thin, and what steps to take before your agent encounters conditions your test suite never anticipated.

The Three Layers of Agent Readiness

"Agent readiness score" means at least three different things in 2026, and each meaning has its own scoring framework, its own tooling, and its own blind spots. Using the wrong checker gives you a number that feels authoritative but measures the wrong thing entirely.

Layer 1: Website and AI Crawler Configuration

Cloudflare's agent readiness tool, launched during Agents Week 2026, scores domains on whether they are correctly configured for AI crawler bots. It checks robots.txt, llms.txt presence, structured data markup, and MCP compliance. GEO Metrics runs a parallel 0-100 score across four dimensions: Discoverability, Content Accessibility, Bot Access Control, and Protocol Discovery, with no account required.

A concrete example illustrates the ceiling problem. A site can score 84/100 on Discoverability by having a clean robots.txt and a valid sitemap, while still serving zero agent-oriented Link headers. The Discoverability score is accurate and honest. It just does not tell you whether an agent can actually negotiate a useful response once it arrives. Partial readiness is the norm, not an edge case.

Layer 2: Codebase and Repository Readiness

Factory.ai released its repo readiness report on January 20, 2026, scoring repositories across 8 technical pillars and 5 maturity levels from Functional to Autonomous. Public scores show real variance: CockroachDB reached Level 4 at 74%, FastAPI landed at Level 3 with 53%, and Express scored Level 2 at 28%. The framework identifies fast feedback loops, clear documentation, and runnable dev environments as the dimensions where most codebases fall short, not model capability.

Layer 3: Organizational and Team Readiness

Skai surveyed 332 practitioners across 9 dimensions and 5 maturity levels. The result: 79% of paid media teams scored below the "Building" tier. More telling is the perception gap. Leaders scored approximately 10 points higher than practitioners on the same assessment. Strategy teams consistently overestimate ground-level readiness, which means organizations may fund agent initiatives before the infrastructure to support them exists.

Where Layer 1 Stops Early

The website layer is where most of the action is, and it is also where the existing scanners stop early. A green crawler-configuration score says an agent can find you and read you. It says nothing about whether the agent can work out who you are (an agent card, an API catalogue, an OpenAPI description), whether it can buy from you (a machine-readable catalogue, a price list it can parse, a payment challenge it can settle), or whether your edge will even let a client with no browser fingerprint through the door. Those are the questions an autonomous buyer asks after discovery succeeds, and they are the ones Moltline's checker was built to answer.

What Moltline's 21-Check Scan Actually Tests

Paste a domain into the free checker at moltlinestudio.com/agent-check.html and it runs twenty-one checks against public URLs, the same ones any crawler can request. No account, no API key, no card; twelve scans an hour per IP. The checks fall into five groups, and every failure comes with the fix. Where the fix is a file that does not exist yet, the checker writes a draft of that file for the domain, with the parts only you can answer spelled out as instructions rather than left blank.

Discovery and identity (6 checks)

Is /llms.txt published? Is /.well-known/agent-card.json present and valid? Is /.well-known/api-catalog served? Is an OpenAPI description reachable? Is /sitemap.xml present and does it parse? Does robots.txt avoid blanket-blocking AI agents? This group overlaps with the crawler-configuration scanners, and it is the part most sites already half-pass.

Machine-readable content (4 checks)

Does the homepage carry JSON-LD structured data, and does that data identify the organisation? Is there a markdown twin of the homepage, so an agent can read the page without burning tokens on markup? Are the title and meta description present and sane?

Commerce and payment (3 checks)

Is a machine-readable catalogue published? Is pricing machine-readable? Is an agent payment challenge discoverable? The last one is the check no crawler-configuration scanner runs: it looks for a payment surface an autonomous agent could actually settle, such as a well-known payment document or an endpoint that answers with a machine-readable HTTP 402.

Trust and security (4 checks)

Does /.well-known/security.txt name a contact? Does HTTP redirect to HTTPS? Is HSTS set? Are the baseline security headers present?

Agent access hygiene (4 checks)

Do the discovery documents send CORS headers? Are JSON documents served as JSON rather than as HTML? Is a client with no User-Agent allowed through, or does a bot wall block it? Does the homepage answer quickly without being enormous?

What the numbers look like in practice

moltlinestudio.com scores 21 of 21 on this scan, and that score is re-measured every day by an automated claims check rather than asserted from memory. For context, Moltline scanned a hundred widely used sites with the same twenty-one checks and published the results as the agent-readiness index; the median site scored 8 of 21. Most of the failures were files that simply did not exist yet.

How to Run the Check

Open moltlinestudio.com/agent-check.html, paste a domain, and read the failures first, because that is the work. Each check reports pass or fail individually with the fix attached, so you see exactly what an agent would trip over rather than a single opaque score. Every scan gets a permanent link that previews with the score, so it reads properly when pasted into Slack or a pull request, plus a badge for a README or footer.

The checker scores a site, not your MCP wiring

The scan reads a public domain. It does not read your repository, your MCP client configuration, or your local environment. To confirm that a hosted MCP server works in your client, paste its Streamable HTTP URL, for example mcp.moltlinestudio.com/<server>, into Claude, Claude Code, Cursor, or Codex CLI, and confirm the expected tools appear in the tool list, then invoke one. Client-side failures such as network policy, client version, or header negotiation are outside what any site scan can observe.

MCPize grades are a separate signal, for servers

If the thing you are evaluating is an MCP server rather than a website, MCPize's audit grades are the external signal to look at. Each Moltline server has a public MCPize result URL that anyone can open without an account. That grade is not part of the site checker's output; the two instruments answer different questions.

Nothing is gated behind the check

The 110 free tools across Moltline's 22 hosted MCP servers are available whether or not you ever run the checker, with no account and no API key. The 50 premium tools unlock with an All-Access licence at $19 per month, billed monthly through NOWPayments and payable in cryptocurrency, or machine-to-machine over x402 at moltlinestudio.com/api. The checker's commerce group tests whether your own site exposes a payment challenge an agent could settle; it is a mirror held up to your domain, not a gate on Moltline's tools.

SKILL.md as a Readiness Signal
SKILL.md as a Readiness Signal

SKILL.md as a Readiness Signal

Site readiness is one question. Whether the skills you hand an agent are well-formed is another, and it matters for the same reason: an agent decides what to do from the metadata it reads at rest. A SKILL.md file is a plain-text skill definition with a small frontmatter block, a name and a description, followed by the instructions the agent follows once the skill activates. Only the metadata is loaded at rest; the full instructions are injected into context when the skill fires. The description doubles as the trigger an agent matches a task against, so a weak description is enough for a model to never use the skill at all. The Agent Skills Explained primer makes the same point from the other direction.

What the linter checks

Moltline runs a SKILL.md linter as a hosted MCP server at mcp.moltlinestudio.com/skillmd-lint, six free tools, no account. Hand lint_skill the full text of a skill file and it checks that the YAML frontmatter is present with the required name and description fields, scores the description's quality, checks the body structure (procedure steps, rules, and what to do when the skill cannot complete), scans for leaked secrets or credentials and for injection-style phrasing, and enforces a size budget. It returns a verdict of pass, pass-with-warnings, or fail, with the errors and warnings listed. A companion tool, packaging_check, validates the file layout of a skill before it is zipped for a marketplace. For reference material, the 138 free SKILL.md files at github.com/GarphenGate/moltline-oss are faster to learn from than documentation: open one, compare its frontmatter and structure against your own, then lint yours.

Lint pass is not sufficient

A file that passes lint but carries a vague description represents a partial readiness failure. The linter scores description quality, but it cannot judge whether the description matches the tasks your users will actually phrase. A skill described as "does the thing" will never fire on a real request. Read the description the way an agent would, as the only thing it knows about the skill before deciding whether to load it.

Why this pairs with site readiness

Site readiness and skill quality are complementary, not redundant. A site can pass all twenty-one checks while the skills your agent carries send it to the wrong tool, and a perfectly written skill cannot help an agent that a bot wall turned away at the door. Neither signal alone is sufficient.

What This Checker Does Not Cover

Moltline's checker tests one thing: whether a public domain is legible, identifiable, purchasable and reachable for an autonomous agent, across the twenty-one checks above. Four boundaries are worth naming directly.

Your own agent project. The checker does not read your repository, your MCP client configuration, or your tool schemas. It cannot tell you whether the tools you wired into Claude or Cursor are correctly configured. That remains a manual test: paste the endpoint, list the tools, call one.

Codebase structure for autonomous coding agents. If your question is whether your repository is structured so an AI coding agent can operate in it autonomously, use Factory.ai's readiness report (account required). It evaluates repositories across 8 technical pillars and 5 maturity levels ranging from Functional to Autonomous. Moltline's checker does not assess repo structure, CI pipeline speed, dev environment runability, or documentation quality.

Organizational and team agentic maturity. If your question is whether your team has the data infrastructure, governance, and processes to run agentic workflows reliably, use Skai's free 3-minute agentic readiness assessment (no signup, 9 dimensions, 5 maturity levels). Moltline's checker does not measure team capability or business process maturity.

Security, uptime, and compliance. The trust group checks four things: a security contact, an HTTPS redirect, HSTS, and baseline headers. That is hygiene, not a security audit. A passing result makes no claim about vulnerabilities, service-level guarantees, or regulatory compliance. No certification is implied.

The overlap is also worth naming. On the discovery checks, Moltline's scan covers similar ground to Cloudflare's tool and GEO Metrics; where it goes further is identity, commerce, and access hygiene. A developer searching for AI agent readiness scoring should leave this page pointed at the right instrument for the layer they are asking about, not carrying a false impression that any one of them covers the others.

x402 and Autonomous Payment Readiness

x402 and Autonomous Payment Readiness
x402 and Autonomous Payment Readiness

The x402 protocol revives the HTTP 402 status code to enable machine-native payments. When an agent requests a gated resource, the server returns a 402 response containing machine-readable payment instructions. The agent reads those instructions, signs a stablecoin transaction, attaches proof in the retry header, and the server grants access. The cycle completes in seconds with zero human approval steps. The /api endpoint at moltlinestudio.com serves a live, inspectable version of this challenge. Issue a raw HTTP request to it and you receive the challenge format in the response; no account, no setup required.

Payment readiness is a distinct readiness dimension from discovery or content. An agent that reaches a paywall it cannot parse stalls at that boundary, and from the outside the stall looks identical to an offline endpoint. Crawler-configuration scanners do not test this condition as a scored dimension, which is why the commerce group exists in Moltline's scan: it checks whether a payment challenge is discoverable on the domain at all, alongside a machine-readable catalogue and price list.

Developers building agents that need to consume paid tools autonomously should provision an on-chain wallet with an appropriate budget scope and confirm the agent can parse a 402 challenge before deployment, not after a silent production stall. Treat it as a pre-launch gate equivalent to environment variable verification. As x402 adoption expands across paid API tooling, any agent deployed without the ability to resolve a payment gate will fail silently on an increasing share of tool calls.

Actionable Takeaways

Before running any checker, identify which layer your question targets. Crawler configuration belongs to Cloudflare or GEO Metrics. Codebase structure for coding agents belongs to Factory.ai. Org and team maturity belongs to Skai. Whether an agent can identify, buy from, and reach your site belongs to Moltline's 21-check scan. Running the wrong checker produces a score that is technically valid but answers a different question than the one you have.

Run the scan on every domain you operate, then on every vendor domain your agents depend on. Fix the missing files first; they are the cheapest failures and the most common.

If you ship skills, lint them with the hosted linter at mcp.moltlinestudio.com/skillmd-lint before they go anywhere near a production agent, and read every description as the trigger it is.

If your agent consumes paid tools autonomously, payment-gate resolution is a production requirement, not a nice-to-have. Hit the /api x402 challenge from your dev environment early and confirm your agent can parse it.

The 110 free tools across Moltline's 22 hosted MCP servers require no signup. Paste mcp.moltlinestudio.com/<server> directly into Claude, Cursor, or Codex CLI. The $19/month All-Access licence unlocks the remaining 50 premium tools, settled in cryptocurrency through NOWPayments, or machine-to-machine over x402 at moltlinestudio.com/api.

Conclusion

Shipping an AI agent without layered readiness testing is not a confidence move; it is a gamble. The core takeaways are clear: readiness is not binary, each architectural layer fails differently, and surface-level end-to-end tests leave massive blind spots in your coverage.

To move forward with intention, start by auditing your current testing strategy against each layer: model behavior, tool integrations, orchestration logic, memory systems, and external dependencies. Identify where your coverage thins out and treat those gaps as production risks, not backlog items.

The teams that ship agents reliably are not the ones with the most tests. They are the ones testing the right things at the right layers. Build your readiness score, know exactly where you stand, and deploy with evidence rather than optimism.

Try it rather than read about it

22 hosted MCP servers, 160 tools, 110 of them free. No account, no API key, no signup — paste a URL into your client and the tools are there.

Browse the servers
← All posts