Donatello AI Benchmark / Protocol & evidence

A method you can inspect—and challenge.

The rules for collecting comparable attempts, keeping failures visible, distinguishing costs and publishing uncertainty. Protocol v1 is published; measured experiments and voting are not yet active.

Published Version 2026-09-11-v1Operated by Donatello

Initial release: protocols and documentation estimates only. No comparative generation results, public votes or quality winners have been published.

Scope: product routes are not interchangeable with model APIs

Keep platform plans and model endpoints separate. Enroll each model with its exact version, provider, access path, plan, safety settings and published price. An undisclosed foundation model is identified as an unknown backend under a dated product route, never attributed to a guessed model. Do not compare API billing with a consumer subscription as if they were identical. This initial release is a protocol plus a documentation study, not a completed experiment.

Prompt and model selection

The first image set has 30 original prompts: three in each of ten creator-oriented categories. They are drafted in plain language before seeing comparative outputs, with task-specific acceptance checks. This is a small fixed English prompt corpus, not a representative sample of all users. Image-reference cases are a separate track. Publish candidate-selection criteria (public access, route documentation, reference support and rights to display results) before enrollment; freeze the enrolled cohort before spending or inspecting outputs. There is no enrolled cohort yet.

Same inputs, visible settings

Use byte-identical prompts and source-reference files, one image per request, three registered repetitions per prompt per model, a 1:1 frame and a 1024 × 1024 target where supported. Freeze all reference bytes and hashes before execution. Prompt enhancement is off where available; unavailable controls and provider defaults are disclosed and stratified rather than silently equated. Equal seeds across different architectures do not make them equivalent. Preserve originals, dimensions and settings, and disclose any separate display normalization. Unsupported configurations are reported, not substituted without a new version.

Attempts, scheduling and stop rules

Publish a cohort manifest and deterministic shuffled round-robin schedule before the first request. The offline planner defaults to a zero-dollar spending authorization and makes no network calls. Three attempts means three registered requests, not three cherry-picked successes. Record refusals, timeouts, downloads that fail and charged errors. A retry is a new attempt; never overwrite the original. Stop all models at the end of a balanced round if the approved cost ceiling is threatened. An incomplete batch cannot establish a winner. Diagnostic extra attempts belong in a separate, clearly labelled run.

Reference ownership and source bias

For face or controlled-voice tests, obtain explicit permission to process and publicly display the source and derivatives. Do not use private customer uploads or public-figure identities. Five of the initial reference cases are blocked pending suitable cleared assets. Three can use a disclosed synthetic vase from our existing public gallery; its Donatello origin is a potential source-domain bias, so it cannot justify general claims about face, character or product-reference superiority. Source diversity must be improved before such claims. Reference fidelity, identity, style and aesthetic quality are reported separately.

Three success definitions, not one

Technical success means the platform returned a successful terminal response. Valid media means the returned bytes decode, match the declared type and meet the registered output constraints. Usable means valid media, every required prompt check satisfied and no severe defect preventing the stated use. A severe defect includes broken essential anatomy, unreadable required text, missing required objects or a wrong product identity. Each criterion gets a separate review decision. At least two reviewers, including one independent of Donatello, must review an output before usability is published; disagreements require a recorded adjudication. These thresholds are editorial operating rules, not a universal definition of image quality.

Blind human evaluation — planned, not active

No votes or human-preference scores exist in this release. Future pairwise voting must compare the same prompt and track, randomize left/right order, conceal provider identifiers until after voting, and offer tie and both-unusable choices. Show the same size and presentation; keep an honest note where a legally required watermark prevents full blinding. Do not strip attribution that a licence requires. Collect preference, prompt adherence, realism and artifacts separately. Donatello staff must not be the only jury, and independent/staff samples must remain distinguishable.

Privacy and voting integrity

Public voting, rate limits and vote storage are not enabled here. Before activating them, review the Privacy Policy and consent/legal basis, define retention and deletion, and minimize identifiers. The planned design uses one vote per approved comparison token, replay protection, engagement checks and server-side abuse limits; aggregate bot-filtering exclusions must be reported. Do not publish voter identifiers, IP addresses, customer prompts, internal request IDs, signed storage URLs or credentials. This phase adds no benchmark-specific visitor or generation telemetry.

Latency and uncertainty

Measure from request dispatch to complete usable media download, including queueing and polling; record execution region and access tier. Summarize successful validated-media latency separately from time spent on failures. Publish sample counts and raw durations. The implemented summary withholds p50 below three valid outputs and p90 below twenty; twenty is a minimum reporting guardrail, not proof of a stable tail estimate. Three repetitions of the same prompt are clustered observations, not independent prompts. Confidence intervals and preference ranks require an adequate independent evaluation sample and a predeclared method; none are claimed now.

Cash, credits and cost per usable output

Keep official list price, actual net charge, credit consumption and required cash commitment separate. Record the source URL and price-verification date, taxes/currency assumptions, refunds, failed-attempt charges, paid tier and promotional credits. A free sample has zero cash paid but can consume scarce credits and time. A subscription's minimum cash commitment is not its marginal API-equivalent cost. Cost per usable output equals total net charges for all registered attempts divided by the number judged usable, only when every charge and required review is complete. With no usable results the finite unit cost is undefined, not $0. Missing prices remain null.

Creator Value: capability, entitlement and completion

The $0/$5/$10/$20 explorer is an explicitly labelled documentation simulation, using a pinned September 11 snapshot. Count image, video, music and voice as four equally counted workflow ingredients only for modality coverage; this is not a quality score. Check the actual selected plan separately. A platform can document all four modalities yet lack the duration, licence or budget needed for a particular project. Completion is measured only from accepted deliverables divided by the fixed brief's deliverables, with modality-level quantities also visible. No full-workflow completion percentage or full cost is inferred from feature lists. Final editing, captions, export rights and additional accounts must be audited in the execution phase.

Commercial use and source hierarchy

Use the official pricing page, Terms, model licence and product documentation applicable on the test date. Record each independently: free access does not establish commercial permission; a paid subscription does not establish clearance for every beta route. Retain unknown and conflicting facts, including free-plan ambiguity. Check permission to publish benchmark inputs and outputs separately from permission to monetize a user's project. Output rights do not guarantee copyright, exclusivity or clearance of third-party likenesses, voices and trademarks. This is a documented permissions review, not individualized legal advice.

Automation, validation and publication gates

The shared types and offline tools support a reproducible schedule, attempt validation, arithmetic checks and byte-level verification of local public media. They do not call paid providers, scrape accounts or publish automatically. Before a measured release: enroll the cohort and budget, capture every attempt, review output safety and rights, hash public files, validate equal coverage and all denominators, audit blinded review, then publish a new immutable data version. Human review remains required for subjective criteria. Syntax and hashes do not by themselves establish that a provider genuinely generated a file.

Versioning and corrections

Each prompt, source snapshot and downloadable dataset is tied to a versioned filename and content checksum. Never silently replace old prompt sets or measured results after inspecting failures. Corrections require a new version and a public explanation, retaining the previous version unless privacy or legal removal is necessary; publish a removal notice in that case. Changing a backend model, source reference, plan or enhancement setting creates a new cohort. The current version contains no measured attempts. Future results will cite their protocol version and retain stable result identifiers; nonexistent results return 404.

Conflict of interest

Donatello.studio operates this benchmark and is a candidate platform under evaluation. This is not an independent laboratory, certification or third-party endorsement. Identical inputs and attempt budgets, full failure retention, prior cohort registration, transparent settings, independent blinded review and separate metrics are intended to reduce bias, not eliminate it. Donatello receives no extra attempts or privileged rubric. Report a rival's better result when supported. No overall utility score or weight selection has been adopted, and no winner is chosen in advance.

Metric dictionary

MetricDefinitionPublication guardrail
Technical success rateSuccessful terminal responses / all attempted requestsNull with no attempts; all refusals and timeouts remain in denominator.
Valid-media rateValidated, conforming outputs / all attempted requestsNull while any successful response is awaiting validation.
Usable-result rateReviewed usable outputs / all attempted requestsNull until validation and required usability reviews are complete.
Cost per usable outputTotal net charges, including failures / reviewed usable outputsNull with missing charges/reviews or zero usable outputs; show why.
p50 / p90 latencyInterpolated quantile of end-to-end validated-media timesMinimum n=3 / n=20 respectively; report n and failure times separately.
Documented modality coverageNumber of confirmed image/video/music/voice capabilities out of fourNot budget entitlement, quality or measured workflow completion.

Methodological references

Reviewed 2026-09-11. These sources informed the design; they do not supply our results or endorse Donatello. The detailed choices above are our own protocol, not a copy of their scoring systems.

Artificial Analysis — image methodology
Separate modalities, documented settings and preference-based evaluation; our sampling choices are specified independently.

Arena — how blind comparisons work
Anonymous pairwise comparison with identity disclosure after a vote; not an endorsement of this project.

Arena — ranking uncertainty
Uncertainty can change the interpretation of apparent ranks. We publish no rank without evaluation evidence.

Google Search Central — Dataset metadata
Dataset metadata describes real downloadable records, not hypothetical benchmark results or promised indexing.