An independent, benchmark-anchored evaluation of how your AI behaves across financial-services regulatory scenarios. Not your tool grading its own homework. We run the evaluation; you receive the report and its stated limitations. Not trust. Proof.
100+ items · 25+ regulatory frameworks · predict-first methodology
scored against a documented gold standard · academic validation in progress
Adoption is racing ahead of the evidence base underneath it. The window to get methodologically-defensible evidence in place is closing.
Regulation (EU) 2026/1744 entered into force on 27 July 2026 and moved the Chapter III high-risk date for Annex III systems to 2 December 2027. Article 50 transparency duties generally began applying on 2 August 2026 — and financial-services governance obligations never waited for either date.
Application dates vary by obligation and system classification. A benchmark can inform governance decisions, but it does not determine legal applicability or certify compliance.
Industry surveys often benchmark adoption; capability benchmarks test financial reasoning; regulatory frameworks specify obligations. The Mergetic Reliability Benchmark connects a defined workflow to documented governance dispositions.
A bad decision doesn't wait for review — it propagates before anyone notices. Three structural problems make today's deployments an unstable equilibrium, and none is solved by the model grading its own work.
When one model both generates a regulated output and judges whether it's acceptable, the judgment is correlated with the generation. When the output is risky, the self-assessment trends with the same risk. The system can't catch its own worst failures.
EU AI Act Articles 12, 14 and 26 presume the recorded rationale can be independently verified. An audit trail that is the model narrating its own output does not satisfy that presumption in the form regulators will require.
Internal evaluations document your own due diligence; vendor benchmarks can be difficult to assess without a transparent method. Mergetic uses a structurally independent evaluation path and a documented reference set.
A productised, done-for-you evaluation. You bring a model or deployment; we run it against a fixed, versioned benchmark and return a scored verdict — with residual risk identified item-by-item, not hidden.
We agree the system under evaluation and the frameworks in scope for your workflow, under a short mutual NDA.
We run the benchmark on our infrastructure against gold-standard answers using a fixed, versioned rubric. Your effort is endpoint access.
You receive the Reliability Evaluation Report — verdict band, dimension scorecard, per-item results, critical findings and recommendation.
A clear basis for whether to deploy, and under what oversight — and a route into governed runtime if you choose it.
The evaluation works whether or not you give us access to your stack — so an IP- or security-sensitive firm is never blocked from getting a verdict.
We send you the benchmark prompts selected for your regulatory regime. You run them through your own AI stack and return the outputs. We score them against the gold-standard answers and produce the report. No access to your systems required.
With endpoint access under NDA, we run the full benchmark against your deployment ourselves, capture the outputs, and score against the gold standard. Your effort is endpoint access and a findings call.
A single cross-regulatory item engaging CBI vulnerable-customer guidance and GDPR Article 22 at once. This depth, times 100+ records, is the benchmark.
CBI CPC 2025 vulnerable-customer guidance and GDPR Article 22 (automated decisions) — two regulatory regimes at once.
Committed before execution: withhold the immediate automated decision and route to human review — on both vulnerability and automated-decision grounds.
Matched the predicted disposition — and surfaced an additional firm-side policy gap the prediction did not anticipate. A deeper finding than expected.
Two-layer scoring — a verdict band first, then a transparent roll-up — so the headline is legible to a risk committee and the reasoning is auditable.
A single Green / Amber / Red band and composite score, readable in seconds.
A is Accuracy, R is Risk and E is Ethical Compliance. S is the composite score—not a fourth assessment dimension.
Side-by-side on identical items: ungoverned behaviour, and behaviour under independent governance.
Every record, colour-coded by verdict. Residual risk surfaced item-by-item, not averaged away.
Fabrications and high-risk failures surfaced, with a deployment-readiness recommendation and oversight conditions.
The scoring method and dimensions, so the verdict can be defended — and reproduced on the next version.
Every item carries a gold-standard answer — the disposition a regulator-aligned expert says is correct, committed before your model runs. Your AI is measured against that, not against its own confidence and not against a generic accuracy metric. That's why an evaluation stands up as third-party evidence where an internal benchmark or a vendor model card won't. Answer maturity is disclosed per item: answers enter at bronze, are promoted to silver on internal cross-validation, and to gold on independent domain-expert review — a validation programme now underway. Every verdict cites the maturity tier it rests on.
The outcome envelope is committed before the item runs, documented in a peer-review-ready paper. Academic benchmarks fix gold answers at publication; predict-first pre-commits the disposition for every record — the difference between a prediction and a rationalisation.
Outputs are scored on weighted dimensions calibrated for the FS risk profile — substantive correctness, citation reliability, disposition appropriateness — not a single accuracy metric.
Distinguishing a wrong output from a fabricated one from an accepted adversarial prompt — not a correct/incorrect binary.
38 items invoke two or more frameworks at once, testing reasoning across overlapping and conflicting requirements — a depth of multi-framework coverage we're not aware of in any public benchmark.
Inter-rater reliability coding of the dataset is engaged with the Applied Innovation Unit, University of Galway, under Innovation Voucher IV20250487.
The dataset has evaluated two independent implementations of the same architecture. The cross-architecture comparison is documented in the in-preparation academic paper.
The market contains evaluation products and runtime controls. Mergetic connects a documented benchmark methodology to a structurally separate runtime governance layer, so evaluation evidence can inform the control path.
Connecting evaluation to runtime governance is an architecture and evidence programme, not a surface feature. It requires a consistent decision model, independent authority and a record that can travel from test to deployment.
And the two products feed each other. Every evaluation sharpens the benchmark; every governed decision generates new scenarios. The asset improves through use.
Judge separated from generator
Per-decision binding veto
Hash-chained, tamper-evident
Any model, any cloud
These are Mergetic design and research claims. Their public status, sources and limitations are available in the Public Claims Register.
The benchmark evaluates how your AI behaves against a documented reference set. Mergetic is the separate, optional next step: a structurally independent governance layer that can issue PERMIT, REGEN or BLOCK before external effect. The two are sold separately and can be adopted in sequence only if you choose. Verify the boundary →
Most AI tools score outputs.
Mergetic gates them.
A controller can approve an output and an independent governance layer — with its own reasoning and its own rulebook — can still block it. That's the structural separation a single model can't reproduce. Fail-closed by design: at the agreed boundary, nothing routed through Mergetic acts without clearance.
Mergetic is designed as an independent runtime governance layer. Illustrative decision records use SHA-256-linked integrity events to demonstrate tamper evidence.
Explore MergeticThe controller approved this output. Governance independently blocked it and routed it to human review — a structural check a single-model pipeline cannot reproduce.
The evaluation reads differently depending on where you sit. Pick the one that fits.
Your FS clients ask you to evaluate their AI against their regulatory obligations — and today you answer with ad-hoc frameworks no regulator would accept. The benchmark's predict-first methodology and cross-regulatory structure aren't replicated by any public competitor. Carry co-branded evaluation reports into client AI-governance files.
Deliverable: co-branded reliability reports, "delivered by [your firm]"FS prospects stall in procurement on "deployment risk" — your model cards read as marketing and audit firms lack a benchmark methodology. A governance-layer reference that complements your model, and an independent harness that gives your sales team a scored report citable in FS RFPs.
Deliverable: scored report on a named model, with a readiness bandYour governance file benefits from methodologically defensible evidence about a deployed AI workflow. A structurally independent, scored evaluation can document a specific regulatory workflow and, where you choose, provide a route into governed runtime.
Deliverable: independent scored evaluation for your AI-governance fileEvery step is low-commitment and reversible — the risk is front-loaded onto us, not you. The deeper material opens up as the conversation gets serious.
This page, the Executive Briefing and the one-page summary. Establishes category and credibility. Forward them freely.
The methodology brief and your sector-specific deck. We identify a candidate evaluation and agree a short mutual NDA.
A sample evaluation against an open model, walked through as a report, so you assess deliverable quality before any commitment.
We run the benchmark against your named system and deliver the report, alongside an engagement letter or LOI.
Mergetic is built and operated by Gratitude Beacons Ltd, an Irish company, anchored in a verifiable IP estate and an independent academic engagement. Every claim on this page has a published artefact behind it.
The architecture is the subject of Irish patent application IE 2025/0516 and corresponding EPO filing EP 26175745.4. The applications are pending and do not imply that a patent has been granted. Verify IP-001 →
Inter-rater validation is in progress with the Applied Innovation Unit, University of Galway, under Enterprise Ireland Innovation Voucher IV20250487. No completed outcome or endorsement is claimed. Verify VAL-001 →
Every factual public statement is bounded by a source, status, limitation and verification date in the machine-readable register. Open the register →
Request a Reliability Benchmark evaluation, or a conversation about where governed runtime could take your deployment.
john@mergetic.com