BENCHLYTIX
  • Leaderboard
  • Methodology
  • Security
  • For enterprise
  • Docs
Check an agent→Sign in
  • Leaderboard
  • Methodology
  • Security
  • For enterprise
  • Docs

Product

  • Leaderboard
  • For developers
  • For enterprise
  • For agents

Trust

  • Scoring methodology
  • Security & verification

Resources

  • Docs
  • Blog
  • Subscribe
  • Changelog
  • Press

Company

  • About
  • Contact
  • Privacy
  • Terms
BENCHLYTIX

© 2026 BenchLytix. Independent AI agent benchmarks.

leaderboard / top agents · this weekLive production data
01attestorcode-generation81Provisional02depguardcode-generation78Provisional
evidence repo materials · sha256-pinned · assessors haiku → sonnet → opus · human review gate · methodology v2.6.0
03agentvet-mcpcode-generation77.5Provisional04mcp-apple-notesmulti-step-reasoning77.5Provisional
Provisional

A desk-assessed score says so. It reads the agent's real repository — pinned to a content hash — and upgrades to Verified only when production telemetry lands. No number pretends.

The problem

Buying an AI agent is a trust decision made on vendor evidence.

01 / DEMOS

The demo is not the product.

Every agent demo is rehearsed on the happy path. Production means malformed inputs, timeouts, retries, and edge cases the demo never met.

02 / STARS

Stars aren't diligence.

GitHub stars measure enthusiasm, not error handling. A starred repository tells you nothing about credential hygiene, scope enforcement, or what happens when a tool call hangs.

03 / SELF-GRADED

Vendors grade their own homework.

Self-reported benchmarks are marketing with a decimal point. An assessment worth trusting is produced by someone who doesn't profit from the answer.

The evidence ladder

The mark is the methodology: three classes of evidence.

Most directories give you one undifferentiated number. BenchLytix labels what kind of evidence produced each score — and only one class can carry an agent to the top.

Analyst estimate

Assembled from public evidence.

Proxy scores for closed enterprise platforms we can't run the pipeline on. Useful for comparison — and always labeled as the weakest class, never dressed up as more.

ceiling 79 · separate proxy rubric
always disclosed as estimate

Provisional

A structured desk assessment.

Three assessor models read the agent's real repository — README, changelog, ecosystem — pinned to a content hash, arbitrated on disagreement, and gated by human review before anything publishes.

ceiling 85 · evidence sha256-pinned
assessor models disclosed per run

Live telemetry

Earned in production. Only here.

Consented, k-anonymized runtime telemetry from the agent's real traffic. It can't be authored, only earned — which is why it's the sole class that lifts a score past the desk ceiling.

ceiling 100 · freshness-gated
lapsed telemetry downgrades honestly

A score can climb the ladder — connect telemetry and Provisional becomes Verified — and it falls back honestly when evidence lapses. The full rules are public: how scores are computed.

The pipeline

How a score is made.

Submittedor indexed from the ecosystem
Evidence fetchedrepo materials · sha256-pinned
Assessed ×3multi-model · arbitrated
Security-scanneddeps · secrets · license
Human-reviewedpromotion gate · auditable
Published + labeledevidence class on every surface

Held out by design. We publish what we measure and how it's weighted — never the exact assessment prompts. A README can't be written to the test, because the test isn't public. Dimensions that saturate get deprecated and replaced.

Two sides of the desk

Built for the buyer. Earned by the builder.

Shortlist on evidence, not enthusiasm.

Compare agents side by side with the evidence class visible on every number. See the security scan, the assessment rationale, and exactly how deep the evidence goes — before anything touches your stack.

  • →Compare any two agentsscores, dimensions, security posture, and evidence class in one view
  • →Read the assessment itselfevery score links to its run — rationale, assessor models, evidence hash
  • →Security scans, disclosed honestlyfailed scans say failed; unscannable agents say so — never "no findings"

compare / any two agents

Side-by-side dimensions, security scans, and provenance labels — with real corpus data, live on the compare surface.

Open the compare tool →
Independence

The score is not for sale.

Independence guaranteeshold for every agent, paying or not
G-01

Scores cannot be purchased

No tier, price, or relationship adds a single point. Paid tiers buy verification depth — a paying agent that assesses poorly publishes a poor score.

✓
G-02

The same pipeline scores everyone

Free, paid, founder-affiliated, and incumbent-proxy agents run the same rubric for their evidence class. No thumb on the scale.

✓
G-03

Every score is attributable

Each run stores the evidence it read — content hash, fetch date — and the assessor models behind the judgment. A score traces to its inputs.

✓
G-04

Evidence class disclosed, never implied

Profile, leaderboard, comparisons, social cards, the embeddable badge — a desk-assessed score is never presented as runtime-verified. Anywhere.

✓

To be precise about scope: BenchLytix is an independent assurance and ratings service for the agent layer — not a regulator, certification authority, or safety auditor. Model-layer certification tells you the model met its thresholds; it tells you nothing about the agent built on top. That gap is what we measure. How we keep scoring independent.

BenchLytix badge for fundry-ddq — a desk-assessed agent — Provisional
a desk-assessed agent — Provisional · live artifact
BenchLytix badge for salesforce-agentforce — an enterprise proxy — Estimate
an enterprise proxy — Estimate · live artifact
The badge

The badge that can't overclaim.

Most badges are stickers. This one is a claim with rules: it says Verified only for live-telemetry agents, labels desk-assessed scores Provisional, and fails safe — if provenance is ever missing, it under-claims rather than over-claims. Embed once; it stays current and stays honest. The examples here are the live endpoints, not mockups.

[![BenchLytix score](https://benchlytix.com/api/v1/badge/your-agent.svg?v=2026-W34)](https://benchlytix.com/badge-click/your-agent)

Top movers, every Tuesday.

Five rising agents. Five droppers. One hidden gem. A weekly pulse on the AI agent leaderboard, free. One-click unsubscribe in every issue.

Subscribe →
Get started

Check the score before you deploy.

Or claim your agent and start earning the only class of evidence that can't be written.

Check an agent's score→Claim your agent
Independent assurance for the agent layer

Every agent looks perfect in the demo.

BenchLytix is the credit score for AI agents and MCP servers — desk-assessed against published evidence, security-scanned, and marked Verified only when live production telemetry earns it. Every score is labeled with the evidence behind it.

Check an agent's score→Claim your agent

118 agents under assessment — every number wears its evidence class

Live telemetryProvisionalAnalyst estimate