The E-ARI Handbook
A complete guide to the platform — what it does, how it works, and why it's built the way it is.
Version 1.1 · September 2026
Foreword
Most AI governance tools are spreadsheets in disguise. They ask a few questions, draw a radar chart, and hand you a PDF that ages out the moment your CFO reads it.
E-ARI was built to do something harder: turn the messy, regulated, fast-moving reality of enterprise AI adoption into a single living record where systems, obligations, evidence and reviewed actions compound over time. One number on a dashboard is easy. A defensible audit trail that survives a Notified Body review three years from now is not.
This document explains every part of the platform — the register that classifies your AI systems and maps the obligations that apply, the evidence and reviewed actions that support them, the exports a reviewer can verify, the monitoring that watches for new gaps, and the readiness assessment and AI agents that diagnose organisational practice.
Read it cover-to-cover or jump to the section you need. There's a glossary at the end if a term feels unfamiliar.
Table of contents
- What E-ARI is
- The problem we're solving
- How E-ARI is structured
- The governance workspace
- Continuous monitoring
- The readiness assessment
- The AI agents
- The user journey
- Methodology & scoring
- Privacy, security, and data handling
- Plans & pricing
- Who this is for
- Roadmap
- Glossary
1. What E-ARI is
E-ARI (the name began as Enterprise AI Readiness Index) is an AI governance and evidence platform built around the EU AI Act. It helps an organisation answer five questions, in order:
| Question | E-ARI answer |
|---|---|
| Which AI systems do we run, and how are they classified? | A register with a rule-decided risk classification per system (Article 5 / 6 / 7 + Annex III). |
| Which obligations apply to each system? | A per-system obligation list, scoped by risk tier and operator role. |
| Do we have the evidence to back it up? | A typed evidence vault with automatic clause extraction, mapping each piece of evidence to the obligations it supports. |
| Who is fixing what, and has it been reviewed? | Improvement actions with an assignee, a due date, linked evidence and a review step. |
| Can a reviewer verify the record? | Sealed exports with signed manifests that an independent verifier can check. |
Alongside the register, an 8-pillar readiness assessment, benchmarked against your sector, diagnoses organisational practice; the weekly compliance digest, the regulatory changelog and reminders before attestation deadlines keep the record current. The output is not a one-shot PDF: systems, obligations, evidence, reviewed actions, FRIAs, technical files and readiness scores stay versioned and exportable.
2. The problem we're solving
2.1 The compliance landscape changed in 2024
The EU AI Act entered into force on 1 August 2024 with phased application, as amended by the Digital Omnibus on AI (Regulation (EU) 2026/1744, in force 27 July 2026):
- 2 February 2025 — Articles 5 (prohibited practices) and 4 (AI literacy) apply.
- 2 August 2025 — Governance, GPAI obligations, penalties, and notified body designations.
- 2 August 2026 — Article 50 transparency duties apply (people must know they are interacting with AI, and AI-generated content must be marked).
- 2 December 2026 — Two further prohibitions (non-consensual intimate imagery; child sexual abuse material) and marking duties for synthetic-content systems already on the market before 2 August 2026.
- 2 December 2027 — High-risk obligations apply for standalone Annex III systems. The Digital Omnibus deferred this date from 2 August 2026 — that is extra build time for a FRIA and an Annex IV technical file, not a reason to wait.
- 2 August 2028 — High-risk obligations for AI embedded in regulated products (Annex I).
Penalties scale to €35M or 7% of global turnover for prohibited-practice violations. There is no "we'll get to it next quarter" version of this.
2.2 Existing tools fail in three predictable ways
| Failure mode | What it looks like | What it costs you |
|---|---|---|
| Spreadsheet maturity models | Four colours on a tab, copied between teams every quarter. | No version control, no defensibility, no link to actual evidence. |
| Big-Four consultancy projects | A 200-page report that's accurate the day it ships. | High cost, low cadence — by the time the next assessment runs, the regulatory landscape has moved. |
| Generic GRC platforms | Compliance modules with a generic "AI" checkbox bolted on. | No native handling of FRIA (Article 27) or Annex IV Technical Files. No risk classifier. No clause-level evidence mapping. |
2.3 What E-ARI optimises for
Three properties that we treat as non-negotiable:
- Defensibility — every recommendation, every classification, every gap is traceable to a clause in a regulation, a sentence in a piece of evidence, or a question in the assessment. No black-box LLM verdicts.
- Compounding — work done in one place feeds the next. Each registered system is classified by rule, and its obligations follow from that classification. Evidence uploaded for one obligation auto-maps to others. Closing a gap on one system updates monitoring across the portfolio.
- Surface area discipline — we deliberately do not try to be a project management tool, a GRC giant, or a model registry. E-ARI sits between your strategy team and your regulators.
3. How E-ARI is structured
The platform is built as four cooperating layers around one record. Each layer has its own data model and UI surfaces, and all of them read and write the same register.
<!-- pdf:diagram-architecture -->| Layer | Name | Components |
|---|---|---|
| 04 | Continuous monitoring | Gap radar · Regulatory changelog · Attestation calendar · Weekly digests |
| 03 | Records and outputs | FRIA generator · Technical file generator · Submission pack · Sealed, verifiable exports |
| 02 | Evidence and actions | Evidence vault · Clause extraction · Controls · Improvement actions with review |
| 01 | AI system register | Registry · Rule-based risk classifier · Applicable obligations |
Alongside the layers, the readiness assessment (8 pillars, 40 questions), pulse checks and the AI agents diagnose organisational practice. Registration does not require an assessment; a system can optionally be linked to a completed one.
4. The governance workspace
This is the core of E-ARI: the record a regulator, a customer or your own internal audit team will ask for. The workspace maps your AI systems to the EU AI Act and produces the artefacts that support each applicable obligation.
4.1 The AI System registry
Every AI system you operate or deploy gets registered with:
- Name & purpose — plain-English description
- Deployer role — provider, deployer, importer, distributor (matters for which obligations apply)
- Sector & populations affected — drives risk classification logic
- Link to your readiness baseline — inherits your organisational context automatically
There is no cap on registered systems on any plan.
4.2 The risk classifier
For each registered system, a classifier runs against the Act's risk taxonomy:
| Risk tier | Trigger | What we generate |
|---|---|---|
| Prohibited (Art. 5) | Subliminal/exploitative practices, social scoring, real-time biometric ID in public spaces | Cease-and-desist guidance + cited articles |
| High-risk (Art. 6 + Annex III) | Safety components, biometrics, education, employment, essential services, law enforcement, etc. | Full Annex IV Technical File + FRIA + post-market monitoring plan |
| Limited risk (Art. 50 / 52 / 53) | Chatbots, deepfakes, emotion recognition, biometric categorisation | Transparency obligation kit (notice templates, disclosure copy) |
| Minimal risk | Everything else | Lightweight obligation list (literacy, voluntary code adherence) |
Each classification ships with a rationale: the specific articles that triggered the tier, the public guidance interpreting them, and the residual ambiguity (so your DPO can override if the regulator's view evolves).
4.3 The evidence vault
A typed object store for compliance evidence:
- Contracts, DPIAs, model cards, dataset documentation, training records, security audits, fairness reports, etc.
- Org-level evidence (applies to all systems) and per-system evidence
- Versioned — replacing a piece of evidence creates a new version, the old one stays linked to historical FRIAs
- Stored in Vercel Blob with private ACL; presigned URLs for download (~5 min TTL)
When you upload, two things happen automatically:
- Classification — an LLM categorises the document type (DPIA, model card, audit report, etc.)
- Clause extraction — the document is chunked, and an LLM identifies which clauses or sentences map to which Article of the AI Act. These mappings are stored as
EvidenceClauserecords and become reusable across obligations.
A DPIA you upload for one system can satisfy parts of three different obligations across two systems automatically.
4.4 The FRIA generator (Article 27)
For high-risk systems, the Fundamental Rights Impact Assessment is mandatory.
E-ARI's generator:
- Pulls your AI system metadata (purpose, populations affected, deployer role)
- Pulls the relevant evidence clauses already mapped to fundamental rights articles
- Generates a structured FRIA covering: process description, time periods, frequency, categories of natural persons affected, specific risks of harm, governance arrangements, complaint mechanisms
- Surfaces gaps inline (e.g. "No evidence found for human oversight design — Article 14 requirement")
- Exports as PDF and
.docxwith embedded source citations
You can finalise a FRIA only when all required sections have backing evidence. Half-finished FRIAs are saved as drafts.
4.5 The Technical File generator (Annex IV)
The Annex IV Technical File is the formal document that a Notified Body or supervisory authority will request first.
E-ARI generates the full structure required by Annex IV:
- General description of the AI system
- Detailed description (development methods, third-party tools, training datasets)
- Information on monitoring, functioning, and control
- Risk management system documentation
- Lifecycle changes log
- Performance metrics, accuracy thresholds, robustness tests
- List of harmonised standards applied
- Declaration of conformity (template)
- Post-market monitoring plan
Each section pulls from your evidence vault and your monitoring plan. Sections without backing evidence are marked clearly so you know exactly what's missing.
4.6 The Gap Radar
A live view of what's missing:
- Per-obligation: total clauses required vs. clauses currently evidenced
- Severity-weighted (a gap on Article 9 risk management ≫ a gap on Article 53 GPAI logging for a non-GPAI system)
- Aging — gaps that have been open for >30 days surface to the top
- Each gap has a one-click "Resolve" workflow that prompts for the specific evidence type needed
4.7 The Submission Pack
A single .zip export, ready to hand to a Notified Body, an auditor, or your own legal team:
- Annex IV Technical File (PDF)
- FRIA (PDF, if high-risk)
- All cited evidence files
- Obligation coverage report
- Monitoring plan
- Manifest with SHA-256 checksums for tamper evidence
4.8 Shadow AI Discovery
Most organisations discover AI tools they did not know were in use: an app-usage export lists them, but nothing reads that list against a risk model. Discovery closes that gap without any integration work: export your app-usage list from Google Workspace, Okta, or Microsoft Entra (or a merchant list from your expense tool), upload the CSV, and E-ARI matches it against a curated catalog of AI tools — each entry annotated with vendor, category, whether the tool trains on customer data by default, and its EU AI Act relevance.
Matched tools appear as undeclared until you register them. One click promotes a discovered tool into the AI System registry (§4.1) with the vendor auto-created in the vendor-risk module — from there the standard pipeline applies: risk classification, obligations, evidence.
4.9 Third-Party AI Vendor Risk
Every third-party AI product is someone else's model and someone else's data terms. The vendor-risk module sends each vendor a versioned ~10-minute questionnaire (data handling, security, AI-specific, legal) through a signed link — no account needed on their side. Scoring is deterministic: every answer carries points, unanswered questions count as worst-case, and critical flags (no DPA, trains on customer data, no deletion path) force a critical rating regardless of the numeric score. Vendor documents (DPAs, SOC 2 reports, subprocessor lists) upload into the evidence vault and run through the same clause-extraction pipeline as system evidence.
4.10 Article 4 AI Literacy Compliance
Article 4 of the AI Act — in force since 2 February 2025 — requires staff to have sufficient AI literacy for their role. E-ARI ships a versioned curriculum (AI fundamentals, EU AI Act essentials, data protection, responsible use), assigned to team members via personal magic links that expire after 30 days. Completions record the quiz score and a SHA-256 attestation hash binding member, module, timestamp, and content version. The whole roster exports as a regulator-ready evidence report; overdue members are nudged automatically each week.
4.11 Continuous Controls
The controls view answers the question an auditor actually asks: which obligations are you meeting right now? Every obligation applicable across your registered systems gets a live state — passing (supported by at least one extracted evidence clause), failing (an open critical/high gap), or pending (applicable but not yet evidenced) — plus warnings for attestations due within 30 days. States are derived from the evidence vault and gap radar, never self-declared. A weekly digest emails you only when something needs action.
4.12 Public API
Growth and above get read access; Enterprise gets write. Endpoints under /api/v1/ expose assessments, the AI registry, vendor-risk results, and derived control states, authenticated with revocable scoped keys. See /developers for the reference.
5. Continuous monitoring
Compliance is not a project. It's a process that runs forever.
5.1 Pulse checks
Lightweight monthly check-ins targeting the questions most sensitive to drift (governance, security, talent). Takes 3–5 minutes. Generates a delta report against your last full assessment.
5.2 Regulatory changelog and compliance monitoring
E-ARI does not scrape a regulatory feed. Regulatory dates come from one curated, versioned source — the regulatory changelog, which publishes the consolidated AI Act text it works from — and every engine reads its dates from that source.
A daily cron job (/api/cron/compliance-monitoring) checks your monitoring plans and sends an attestation-due reminder when a system's next attestation falls within the next 30 days — at most one reminder email per system per week, plus an in-app notification. The weekly digest (§5.4) covers failing controls, pending evidence, and attestations due.
5.3 Attestation calendar
Annex IV requires periodic attestations for high-risk systems. The platform tracks:
- Annual conformity reassessment due dates
- DPIA refresh windows (every 24 months for evolving systems)
- Post-market monitoring report deadlines
- Internal review cycles (configurable per system)
When a deadline falls within the next 30 days, the daily compliance cron emails an attestation-due reminder — at most one per system per week — and creates an in-app notification. That is the shipped reminder behaviour; there is no tiered 60/30/7/1-day cadence.
5.4 Email & in-app notifications
All notifications are delivered via the unified template system:
- Weekly compliance digest (paid tiers — sent only when something needs action)
- Scheduled drift detection with alerts on change — pulse checks and the monitoring engine flag score regression; alerts fire on change, not continuously
Unsubscribe and preference management at /portal/preferences.
6. The readiness assessment
The readiness assessment is a supporting diagnostic of organisational practice. Registering systems does not require it, and neither its score nor its readiness level is a compliance determination.
6.1 The 8 pillars
We score readiness across eight pillars, with weights derived from research literature on AI adoption success factors. The weights sum to 1.00.
| # | Pillar | Weight | What it measures |
|---|---|---|---|
| 1 | Strategy & Vision | 15% | Formal AI strategy, executive sponsorship, ROI measurement, multi-year roadmap. |
| 2 | Data & Infrastructure | 15% | Data quality, accessibility, governance, ETL/MLOps maturity, compute. |
| 3 | Technology & Tools | 12% | ML platforms, model lifecycle tooling, vendor diversity, cloud architecture. |
| 4 | Talent & Skills | 13% | Hiring pipeline, internal training, role definitions, succession planning. |
| 5 | Governance & Ethics | 15% | Ethics review boards, policy completeness, bias auditing, accountability. |
| 6 | Culture & Change | 10% | Adoption psychology, change management, communication, resistance handling. |
| 7 | Process & Operations | 10% | Workflow integration, process re-engineering, feedback loops, SLA management. |
| 8 | Security & Compliance | 10% | Model security, privacy controls, regulatory alignment, audit readiness. |
Each pillar contains exactly 5 questions on a Likert scale (1–5). 40 questions total. Median completion time: ~15 minutes.
6.2 Maturity bands
Once weighted, the overall score (0–100) maps to one of four bands:
| Band | Range | Interpretation |
|---|---|---|
| Laggard | 0–25 | Minimal foundational elements. AI initiatives are ad-hoc or absent. |
| Follower | 26–50 | Early-stage readiness. Some initiatives but lacking cohesion or alignment. |
| Chaser | 51–75 | Progressing readiness. Foundations in place, active investment underway. |
| Pacesetter | 76–100 | Advanced readiness. Well-positioned for AI-driven competitive advantage. |
Bands are deliberately broader than a percentile — they're meant to drive behaviour (where to invest), not vanity metrics.
6.3 What you get back
A completed assessment produces a structured Assessment record:
- Overall score (0–100, sector-weighted) and baseline score (unweighted) for transparency
- Per-pillar scores with question-level breakdowns and applied interdependency adjustments
- Maturity band with a plain-English interpretation
- X-Ray findings — structural risk patterns detected from response combinations (see §9.4), each with severity, evidence trail, business impact, and a concrete next move
- Sector weighting application — which pillars were emphasised for your sector and why
- Sector benchmark comparison (your scores vs. curated industry averages)
- Strength / weakness matrix — top three pillars and bottom three
- Initial regulatory mapping — pillars that align with applicable regulations (GDPR, EU AI Act, ISO 42001, NIST AI RMF)
- Leverage analysis — the pipeline is re-run with each answer improved one step, ranking every possible move by its exact overall-score gain (including interdependency-rule releases and sector weighting), plus a greedy simulated path to the next maturity band
This record is the seed for everything else. Crucially, the X-Ray findings are what every downstream agent grounds its narrative in — so the report you receive is anchored in your specific structural patterns, not a generic write-up of your numbers.
7. The AI agents
When an assessment finishes, an orchestrator runs a six-agent pipeline. Each agent has a single job, and the outputs of earlier agents feed the inputs of later ones.
<!-- pdf:diagram-agents -->The pipeline executes in four sequential stages:
| Stage | Agents | Mode | Role |
|---|---|---|---|
| 1 | Scoring | Sequential · runs first | Deterministic math — produces the numerical baseline |
| 2 | Insight + Discovery | Parallel | Run simultaneously since their inputs don't overlap |
| 3 | Report | Sequential · compiles | Combines outputs from stages 1–2 into the executive document |
| 4 | Literacy | Sequential · adaptive | Personalised curriculum based on weakest pillars |
| — | Assistant | On-demand | Interactive chat, available post-pipeline |
7.1 Scoring Agent
Deterministic, no LLM involvement, runs first. Executes an eight-step pipeline:
- Validate — every required question is answered, a Don't-know or Not-applicable annotation counting as answered; a pillar with fewer than three of its five questions rated is refused (
E_INSUFFICIENT_PILLAR_COVERAGE), and one with none at all is refused asE_NO_PILLAR_COVERAGE - Normalize — each pillar mapped to 0–100 from its Likert sum, over the questions actually rated: the bounds follow the answered count, so an annotated question neither drags the pillar down nor pads it up
- Adjust — six documented cross-pillar interdependency rules fire (§9.2)
- Baseline composite — unweighted overall preserved for transparency
- Sector weighting — pillar weights re-balanced for the organisation's sector and renormalised to 1.0 (§9.3)
- Composite & classify — final overall score, maturity band, critical-failure flags
- X-Ray detection — eight structural pattern detectors scan response combinations (§9.4). A detector whose declared inputs include an annotated question does not run and is reported as "could not be evaluated", never as one that ran and found nothing
- Leverage simulation — the entire pipeline is re-run with each answer raised one step, producing an exact score-gain ranking of every possible improvement and a greedy simulated path to the next maturity band. Because the engine is deterministic, these are reproducible facts, not estimates — the same property that lets a client audit their score lets them audit their roadmap
The scoring math is open and reproducible. It is the only deterministic step in the pipeline — every other agent grounds its outputs against this scoring result and its X-Ray findings, which means LLM hallucinations cannot move your score, and the narrative cannot drift away from your actual structural patterns.
7.2 Insight Agent
LLM-powered, generates narrative insights grounded in the scoring result. For each pillar:
- A 2–3 sentence interpretation of the score in business language
- Two specific risks if the score stays where it is
- Two specific opportunities unlocked by reaching the next maturity band
Outputs are run through a grounding verifier — claims that don't tie back to a numeric score in the scoring result are stripped out before display.
7.3 Discovery Agent
Runs in parallel with Insight because their inputs don't overlap. Discovery answers the question "what does your organisation actually look like to the outside world?":
- Auto-classifies your sector via a 30-token LLM call (no cross-industry pollution)
- Pulls public signals via Tavily Search and Extract (recent news, regulatory filings, public hiring patterns)
- Synthesises a contextual brief: sector trends, peer announcements, observed regulatory pressure
- Citations are inline
[S1],[S2]— no source, no claim
If the public signal is weak or untrustworthy, Discovery declines to synthesise rather than fabricating a "medium-confidence" reading.
7.4 Report Agent
Compiles everything above into a structured executive document:
- Executive summary (one page, weighted by your sector context)
- Pillar-by-pillar deep dive with insights and risks
- Top three recommended actions, each tied to specific pillars and obligations
- Benchmarking section with the sector context from Discovery
- Appendix: full question responses, methodology notes, regulatory mapping
Available as an in-app view, a .docx download, and a print-optimised PDF.
7.5 Literacy Agent
Adaptive, runs after Report. Looks at your weakest pillars and assembles a personalised mini-curriculum:
- Targeted articles from the Literacy Hub
- Quizzes calibrated to your maturity band
- Recommended learning paths (e.g. "From Follower to Chaser on Governance" — 4 modules, ~3 hours)
Unlike a static knowledge base, Literacy never recommends content for pillars you're already strong in. It earns its position in the pipeline by being personalised.
7.6 Assistant Agent
On-demand, post-pipeline. A grounded chat interface that has read-only access to your assessment, your AI systems, your evidence, and your obligation coverage.
- Asks: "What's blocking us from being a Chaser?"
- Refuses to answer questions outside its data scope (no general AI advice — only what's grounded in your platform state)
- All citations link back to specific records in your account
8. The user journey
Your portal overview follows the same four steps, with the next action shown first.
Step 1 — Register
Start with an AI system you actually use; the Shadow AI scan finds tools nobody declared. The rule-based classifier sets each system's risk tier and the obligations that apply. Registration does not require a readiness assessment.
Step 2 — Govern
Review classifications, link evidence to the obligations it supports and resolve the improvement actions tracked against your systems. Controls show each obligation as evidenced, failing or awaiting evidence.
Step 3 — Assess
Answer the 40-question readiness assessment for a score and maturity band across 8 pillars — a diagnostic of organisational practice whose priorities become improvement actions.
Step 4 — Maintain
Review systems, evidence and readiness regularly. Follow the regulatory changelog and the attestation calendar, export sealed records a reviewer can verify, and run pulse checks between assessments.
9. Methodology & scoring
9.1 Pillar normalisation
For each pillar p with questions q₁ ... q₅ answered on a 1–5 Likert scale:
pillar_score(p) = (Σ qᵢ - 5) / 20 × 100
This normalises the raw 5–25 sum to a 0–100 scale.
9.2 Interdependency rules
Six documented rules fire on the normalised pillar scores. Each rule has a published rationale and is logged on the assessment record so any adjustment can be reproduced.
| Rule | Trigger | Effect | Why |
|---|---|---|---|
| R1 | Governance < 30 | Technology × 0.70 | Mature tooling without governance is uncontrolled risk, not capability. |
| R2 | Data < 30 | Strategy × 0.85 | Strategy without a data foundation is aspirational, not operational. |
| R3 | Security < 30 | Technology × 0.85 | Deployable models without controls are a liability. |
| R4 | Talent < 25 | Strategy × 0.90 | Ambition that exceeds internal capacity does not ship. |
| R5 | Governance < 35 and Security < 35 | Process × 0.85 | Industrialised AI without controls compounds incidents at scale. |
| R6 | Culture < 30 | Process × 0.92 | Change-resistant cultures cannot operationalise AI gains. |
Rules apply in declared order. The original score, the adjusted score, and the rule that fired are all persisted on the assessment for audit replay.
Compounding: rules apply in order to the running value, so each rule acts on whatever the previous rule left behind. Worked example: Governance 25, Security 25, Technology 75: R1 gives 52.50, then R3 applies to 52.50 giving 44.63. Had R3 acted on the original 75 instead of the running value, it would have produced 63.75 — the audit log records each rule's input and output, which is exactly what makes the chain replayable.
9.3 Sector weighting
Default pillar weights live in the engine (15% Strategy, 15% Data, 12% Technology, 13% Talent, 15% Governance, 10% Culture, 10% Process, 10% Security). After interdependency adjustments, sector-specific multipliers are applied to those weights and the result is renormalised to sum to 1.00. This keeps the overall score on the 0–100 scale while reflecting what actually matters in each sector.
| Sector | Pillars emphasised (×) | Pillars de-emphasised (×) |
|---|---|---|
| Healthcare | Governance ×1.35 · Security ×1.25 · Data ×1.20 | Technology ×0.90 · Culture ×0.90 |
| Finance | Governance ×1.30 · Security ×1.25 · Data ×1.20 | Culture ×0.85 · Talent ×0.90 |
| Manufacturing | Process ×1.30 · Data ×1.20 · Technology ×1.15 | Culture ×0.85 · Talent ×0.90 |
| Retail | Data ×1.30 · Process ×1.20 · Technology ×1.10 | Governance ×0.90 · Talent ×0.95 |
| Technology | Technology ×1.25 · Talent ×1.20 · Strategy ×1.10 | Governance ×0.90 · Security ×0.95 |
| Government | Governance ×1.35 · Security ×1.25 · Data ×1.10 | Talent ×0.90 · Culture ×0.90 |
| Energy | Security ×1.30 · Data ×1.20 · Process ×1.15 | Talent ×0.90 · Culture ×0.90 |
| Education | Governance ×1.25 · Culture ×1.20 · Security ×1.10 | Process ×0.95 · Technology ×0.95 |
The unweighted baseline overall score is preserved alongside the sector-weighted overall, so reports can show how sector context moved the number — and why.
9.4 X-Ray engine — structural pattern detection
A score is a number. A pattern is a story. Two organisations can both land at 47% — one because they're evenly mediocre, one because they have aggressive AI ambition with no data foundation. The interventions for those two are completely different.
The X-Ray engine runs eight deterministic detectors over the response map and surfaces structural failure modes that emerge from how questions interact across pillars. Each finding carries severity (low / medium / high / critical), the question-level evidence that triggered it, the business impact in plain terms, and a single concrete recommended move.
| ID | Pattern | Trigger | Why it matters |
|---|---|---|---|
| P-01 | Shadow IT Risk | High technology adoption + low governance floor | The single fastest path to an EU AI Act non-compliance finding. |
| P-02 | The Ambition Gap | High strategy ambition + broken data/governance foundation | Flagship initiatives stall in pilot at months 9–12. |
| P-03 | Pilot Purgatory | High experimentation appetite + low MLOps maturity | Pilots stall at the demo when tooling exists without the process to operationalise it — and appetite for the next attempt falls with each one that stalls. |
| P-04 | Compliance Cliff | (governance_5 + governance_3 + security_5 ≤ 6) AND (technology_3 ≥ 4 ∨ process_4 ≥ 4 ∨ technology_3 + process_4 ≥ 2) | Article 4 literacy binds today whatever you deploy; standalone Annex III high-risk obligations apply from 2 December 2027, and a FRIA plus technical file take quarters to stand up. Critical when real capability is being deployed into that gap. |
| P-05 | Talent-Strategy Mismatch | High AI ambition + low internal talent supply | Roadmaps slip and external spend grows to cover the gap; each missed date buys another contractor instead of the internal capability that would have prevented it. |
| P-06 | Bias Blindspot | High deployment maturity + low fairness/transparency | A bias incident in this shape is discovered externally rather than internally, because nothing is monitoring for it — the same problem reported by a customer or regulator reads as an oversight failure under Article 14. |
| P-07 | Executive Disconnect | Strong operational capability + low executive sponsorship | Capability without sponsorship fragments; AI engineers leave inside 18 months. |
| P-08 | Process-First Mismatch | High process maturity + weak data layer | Disciplined teams industrialise the wrong outputs at scale. |
The X-Ray block is then prepended to every downstream agent's prompt. The Insight Agent must map every risk and every next-step it produces back to a finding ID. The Roadmap Agent must neutralise every CRITICAL finding in Phase 1 and every HIGH finding in Phase 2. The Discovery Agent must reflect the findings in its gap indicators. This is what makes the report feel tailored — the agents are grounded in your specific pattern set, not paraphrasing a number.
9.5 Confidence interval
A scoring result includes a confidence band derived from:
- Question completion rate — partial assessments get wider bands
- Response variance — extreme highs and lows in adjacent pillars suggest survey fatigue
- Free-text consistency — when an elaboration contradicts the Likert score, confidence drops
The band is shown in the UI as a range (e.g. "68 ± 4") and is preserved in exports.
9.6 Sector benchmarks
Benchmarks are curated from public research (industry association reports, academic studies, regulator-published statistics). They are explicitly labelled in the UI:
Sector benchmarks are AI-estimated and intended for directional guidance only. Actual sector performance may vary.
We never claim a benchmark is current — we cite the year of the underlying data.
9.7 Why we don't use a single "trust score"
Some platforms compress everything into a single 0–100 "trustworthiness" number. We don't. A single number obscures the only useful signal: which dimension is failing. An organisation can be a Pacesetter on Strategy and a Laggard on Security simultaneously — that's actionable. Averaging it to "Chaser, 62" is not.
Our methodology, honestly stated
E-ARI asks an organization to rate its own AI readiness. The published literature on self-assessment is unambiguous about what that can and cannot measure, and this page states both. Every design decision below names its source.
The three-tier model
- Self-reported — The questionnaire as answered — a perception measure, labelled as one.
- Evidence-backed — Answers with artifacts in the vault — verified existence, a lower bound on actual readiness.
- Independently-verified — Answers verified by an auditor or by the deterministic engine against the evidence vault — verified operation. Not yet offered.
A completed assessment therefore surfaces three numbers, not one: the self-reported score (labelled "Self-reported — a perception measure"), the evidence-backed score (labelled "Evidence-backed — only answers with vault artifacts"), and the coverage share ("X% of answers are evidence-backed"). When no answers carry evidence, the evidence-backed score is not shown at all — a self-reported number with nothing behind it is not a second opinion, it is the same opinion.
Answers without evidence are counted at half weight in the evidence-backed score. The 0.5 is our choice, not a constant from the literature. What the research establishes is the direction: self-ratings are lenient by about a third of a standard deviation (d = 0.32; Heidemeier & Moser, 2009), and acquiescence pushes agreement-scale answers up (Krosnick, 1999). Neither hands us a number. We halve what nobody can check, round on purpose so it is not mistaken for a measurement, and report the coverage share beside it so you can see how much of the score the weight is acting on at all. Discounting perception — not deleting it — is what the tier model claims.
The multi-rater design
A single respondent answering all 40 questions is a single-informant design. Conway & Huffcutt (1997) put inter-rater reliability at .50 for supervisors, .37 for peers and .30 for subordinates — and agreement BETWEEN sources lower still, .22 for self versus supervisor. Even the best case leaves half the variance with the rater rather than the organization. Multi-rater mode invites up to four functional leaders — CTO for Technology & Data, DPO for Governance & Security, COO for Process & Culture, CHRO for Talent & Strategy — to answer only the pillars their role covers. The consensus score is the median answer per question across raters, scored by the same deterministic engine. Single-rater assessments are flagged as single-rater, with that caveat attached.
Divergence between raters is a signal, not an error — it indicates where functional leaders have different views of the organization’s readiness.
Don’t-know, not-applicable, and the respondent role
A forced-choice agreement scale gives a respondent who genuinely does not know only wrong options. The honest ones pick "Strongly Disagree" — and the instrument records a finding the respondent never made: a governance deficit that is actually a knowledge gap. Schuman & Presser’s (1981) no-opinion-filter experiments are the classic result — forced-choice forms collect pseudo-opinions from people who hold none — cited for direction only: that source is NOT INDEPENDENTLY VERIFIED, and no figure from it is used. The assessment therefore records two annotations an answer can carry instead of a rating — Don’t know, and Not applicable — and both are excluded from the scoring arithmetic: the pillar renormalises over its answered questions, so an unknown never drags a score down and never pads one up. Both are surfaced as coverage gaps instead — per pillar, "x of 5 answered", with the unanswerable questions named. A pillar every question of which is annotated is refused outright (E_NO_PILLAR_COVERAGE): a score computed from no measurements is not a score. Renormalising is bounded for the same reason — it borrows the answered questions to stand in for the missing ones — so a pillar with fewer than three of its five questions rated is refused too (E_INSUFFICIENT_PILLAR_COVERAGE) rather than let two answers stand for a pillar. Unbounded, the arithmetic reached its limit at one rated answer per pillar: eight of forty, all at the top of the scale, produced a composite of 99.99 that looked exactly like one computed from the whole instrument. The X-Ray detectors do not fire on an annotated question, for the same reason as the refusals — their trigger cannot be evaluated on data that does not exist — and a detector that could not be evaluated is reported as such, never as one that ran and found nothing. This is a scoring change, not a UI addition: an assessment that uses the annotations scores differently from one where the same questions are answered 1. The annotations shipped as SCORING_VERSION 5.5 and the proration bound as SCORING_VERSION 5.6, each with re-derived conformance vectors; the bound moves no score, and assessments scored without annotations remain bit-identical to 5.4.
"Single-rater" is the blunt version of a sharper question: which single rater? A CTO answering all 40 questions speaks first-hand to Technology & Data and second-hand to Governance & Culture, and the reliability cost is not spread evenly (Conway & Huffcutt, 1997; the 360° practice of matching raters to domains is Bracken, Timmreck & Church, 2001). Every assessment therefore records the respondent’s functional role, and the results name which pillars that role covers first-hand — the second-hand pillars are labelled as such, so a reader can see where the perception is thinnest. The role also seeds multi-rater invitations: the roles the respondent does not hold are the ones worth inviting.
The second instrument: regulatory constructs
The redesign audit also asked what regulatory ground the 40 questions do not cover at all, and found nine constructs: an AI-system inventory with recorded Art. 6 classifications, vendor and GPAI supply-chain duties, incident management, logging, the specifics of human oversight, post-market monitoring, the depth of bias metrics, and the ISO/IEC 42001 internal-audit (9.2) and management-review (9.3) loops. They are deliberately not added to the scored instrument — the 40 questions are frozen methodology (SCORING_VERSION 5.6), and appending questions would rewrite every assessment ever taken. They form a second instrument instead: a diagnostic checklist answered yes / no / partially / don’t-know / not-applicable, where no and partially are open gaps, don’t-know is a knowledge gap, and not-applicable is recorded and excluded. It is reported as diagnostics and gaps — never as a score. Which obligations apply is decided by the compliance engine; the items cross-link its obligation codes to show which duty a gap puts pressure on, nothing more.
Behaviourally-anchored verification
Where a question has a factually-checkable form ("Does your incident-response plan specifically address AI system failures?" rather than "How mature is your incident response?"), the assessment offers both: the original Likert item, which stays the scored answer for compatibility, and a "Verify this answer" control with a Yes/No/Partially answer. A verified Yes counts at full weight in the evidence-backed score; the Likert answer is unchanged. This is the BARS pattern: asking whether a specific, observable thing is true rather than how good it feels. Latham, Wexley & Pursell (1975) and Schwarz (1999) support the mechanism — the format sets the reference frame, and a checkable question is harder to inflate than a global one. Neither reports a percentage reduction, and the "~40%" this document previously claimed was not in either.
The bias research this is grounded in
- Podsakoff, MacKenzie & Podsakoff (2012). "Sources of Method Bias in Social Science Research and Recommendations on How to Control It." Annual Review of Psychology 63, 539–569. — Common-method variance — one person answering both sides — inflates observed relationships on subjective constructs. The review establishes the direction and how to design against it; it does not give a single percentage, and the earlier "25–40%" here was unsourced.
- Kruger & Dunning (1999). "Unskilled and Unaware of It." Journal of Personality and Social Psychology 77(6), 1121–1134. — Bottom-quartile performers scored at the 12th percentile and placed themselves at the 62nd — 50 points, in the paper’s own words. The orgs least equipped to govern AI are the likeliest to rate themselves highly, because they cannot see what mature looks like.
- Heidemeier & Moser (2009). "Self–Other Agreement in Job Performance Ratings: A Meta-Analytic Test of a Process Model." Journal of Applied Psychology 94(2), 353–370. — Across 128 samples, self-supervisor agreement is r = .22 and self-ratings are lenient by d = 0.32 (Δ = .49 corrected) — about a third of a standard deviation, not the half previously stated here.
- Latham, Wexley & Pursell (1975). "Training Managers to Minimize Rating Errors in the Observation of Behavior." Journal of Applied Psychology 60(5), 550–555. — Raters asked about specific observable behaviour err less than raters asked for a global judgment. This is a rater-training study, not a BARS-versus-Likert comparison, and it reports no percentage — the "~40%" previously stated here was not in it.
- Schwarz (1999). "Self-Reports: How the Questions Shape the Answers." American Psychologist 54(2), 93–105. — Question format sets the reference frame: "How mature is your incident response?" invites a global judgment; "Does your incident-response plan address AI system failures?" invokes a checkable fact.
- Krosnick (1999). "Survey Research." Annual Review of Psychology 50. — Acquiescence — agreeing regardless of content — is a known threat to agreement-scale items, reduced when questions ask for facts rather than endorsement. No percentage is attributed here: the earlier "~10–15%" was unsourced, and a figure of that size would not in any case justify halving an answer. The 0.5 discount is our own conservative choice, described below.
- Conway & Huffcutt (1997). "Psychometric Properties of Multisource Performance Ratings: A Meta-Analysis of Subordinate, Supervisor, Peer, and Self-Ratings." Human Performance 10(4), 331–360. — Reliability differs by source: supervisors .50, peers .37, subordinates .30 — the .50 is the best case, not the average. Agreement between sources is lower still (self-supervisor .22). Any single rater leaves most of the variance rater-specific, which is why multi-rater mode and the single-rater flag exist. The gain from averaging raters is Spearman-Brown, our projection rather than a finding of this paper.
- Van der Heijden & Nijhof (2004). "The value of subjectivity: problems and prospects for 360-degree appraisal systems." International Journal of Human Resource Management 15(3), 493–511. — Divergence between raters from different functional backgrounds is itself a valid diagnostic signal — the basis for reporting divergence instead of averaging it away.
- Bracken, Timmreck & Church (2001). The Handbook of Multisource Feedback. Jossey-Bass. — The established 360° practice: raters answer question sets matched to their domain expertise, and disagreement is reported as a finding.
- Oxford Insights (2024). Government AI Readiness Index Methodology. — Uses externally verifiable data sources wherever possible and discloses survey data as "perception-based" — the honesty pattern the tier labels follow.
- Cisco (2024). AI Readiness Index Methodology. — States plainly: "This is a perception study reflecting business leaders’ views" — a self-assessment labelled as perception, not objective readiness.
What we can claim
- "The self-reported score is a perception measure — we label it as such"
- "The evidence-backed score counts only answers with artifacts in the vault — it’s a lower bound on actual readiness"
- "The coverage share tells you how much of the self-assessment is backed by verifiable evidence"
- "Multi-rater mode surfaces disagreement between functional leaders as a diagnostic signal"
- "Don’t-know and not-applicable answers are excluded from the score and reported as coverage gaps — never converted into the worst answer"
- "Every assessment records the respondent’s functional role, and the results name which pillars that role speaks to first-hand and which it does not"
- "The constructs instrument maps nine regulatory areas the scored instrument does not cover, reported as diagnostics and gaps — never as a score"
What we still cannot claim
- "A higher E-ARI score predicts AI project success — no outcome data"
- "The instrument is independently validated — no published validation study (this requires 50+ orgs and 12 months)"
- "The evidence-backed score equals audited compliance — it verifies existence, not operating effectiveness"
The validation study
Data collection for a validation study has started. Consent-gated organisations get one monthly snapshot of their scores — self-reported, evidence-backed, coverage share, pillar map, scoring version — and the outcome side is recorded alongside: self-reported project, incident and audit events, plus the compliance countdown’s own accuracy ledger, joined to the scores only at analysis time. Withdrawal deletes an organisation’s snapshots and outcomes. None of this changes a claim: no validity number exists, and the "what we still cannot claim" list above is unchanged until a study on this data is published. Study readiness is published as a counter — organisations and months — wherever this methodology is disclosed, so the distance still to travel is stated rather than implied.
These limits are load-bearing. When the instrument has outcome data, a published validation study, or auditor sign-off behind a tier, the relevant line moves from "cannot" to "can" — and not before.
10. Privacy, security, and data handling
10.1 What we store
| Category | Examples | Where |
|---|---|---|
| Account info | Name, email, organisation, role, password hash | PostgreSQL (Supabase) |
| Assessment responses | Likert answers, free-text elaborations | PostgreSQL |
| AI system metadata | System name, purpose, sector, classification | PostgreSQL |
| Evidence documents | Uploaded files (DPIAs, contracts, model cards) | Vercel Blob (private) |
| Generated artefacts | FRIAs, Technical Files, reports | PostgreSQL + Vercel Blob |
| Logs | Compliance audit log, user actions | PostgreSQL with 90-day retention |
| Payment | Customer ID, subscription tier | Stripe (we never store card numbers) |
10.2 Encryption
- In transit — TLS 1.3 everywhere, HSTS preload-listed
- At rest — AES-256 (provided by Supabase and Vercel Blob)
- Passwords — bcrypt with cost factor 12
10.3 Authentication
- Email + password (NextAuth credentials provider)
- Google OAuth (NextAuth Google provider)
- SSO (OIDC) on Enterprise tier — configured with your identity provider (Okta, Entra ID, Google Workspace, Keycloak). When enforcement is on, password sign-in is refused while the identity provider is registered; if it is not registered, password sign-in stays available as break-glass access.
Sessions use JWT with HttpOnly, Secure, SameSite=Lax cookies.
10.4 Data residency
Compute is pinned to an EU region (Vercel, Frankfurt). The database runs on Supabase (AWS eu-central-2, Zurich); Switzerland is outside the EU but covered by a European Commission adequacy decision. The default model provider processes data in China and other sub-processors operate in their own regions (see /security#subprocessors), so we do not claim EU-only data residency. Cross-border transfer safeguards are described in the DPA.
10.5 Data deletion
- Account deletion is self-service from
/portal/preferences - Deletion is hard (not soft) — within 30 days, all account data including evidence files, assessments, FRIAs, and logs are purged
- Audit log retention is independent and purged after 90 days regardless of account state
10.6 GDPR posture
- DPA published at
/data-processing - Sub-processor register maintained at
/security#subprocessors - Right to erasure honoured within 30 days
- Right to portability — full account export as JSON from
/portal/preferences
10.7 Third-party LLM use
The platform's default model provider is DeepSeek, which processes data in China; China has no European Commission adequacy decision, so transfer safeguards must be assessed. Model-assisted features are narrative generation, clause extraction from uploaded evidence, compliance drafting and organisation enrichment; readiness scoring and rule-based risk classification use no language model. Retention and training conditions follow the provider's API terms: E-ARI claims no negotiated zero-retention or no-training commitment, and trains no models of its own on customer data. Evidence text sent for clause extraction passes through regex PII redaction first.
11. Plans & pricing
| Plan | Price | Best for |
|---|---|---|
| Starter | Free | Registering systems and seeing the obligations that apply |
| Professional | €199/month or €1,990/year (save 17%) | A compliance lead linking evidence and producing verifiable exports |
| Growth | €499/month or €4,990/year (save 17%) | Governance teams of up to 250 people |
| Autopilot | €15,000/year (€1,250/month, billed annually) | Mid-market continuous compliance, sales-assisted |
| Enterprise | Custom | Regulated industries that need SSO and write access to the API |
11.1 What's included by tier
| Capability | Starter | Professional | Growth | Autopilot | Enterprise |
|---|---|---|---|---|---|
| AI system register + rule-based classification | ✓ | ✓ | ✓ | ✓ | ✓ |
| Improvement actions with review | ✓ | ✓ | ✓ | ✓ | ✓ |
| Evidence records + sealed, verifiable exports | — | ✓ | ✓ | ✓ | ✓ |
| Org-wide Article 4 literacy report | — | — | ✓ | ✓ | ✓ |
| FRIA, technical file, gap analysis + submission pack | — | — | — | ✓ | ✓ |
| Shadow AI discovery scans / month | — | — | 1 | Unlimited | Unlimited |
| Vendor risk questionnaires | — | 5 | 25 | Unlimited | Unlimited |
| Team members (Article 4 roster) | 5 | 50 | 250 | Unlimited | Unlimited |
| Full assessments / month | 1 | 5 | 20 | Unlimited | Unlimited |
| Pulse checks / month | 3 | 15 | 50 | Unlimited | Unlimited |
.docx reports | 1/mo (extra exports on request) | 3 included | Unlimited | Unlimited | Unlimited + custom branding |
| Literacy Hub | Basic | Full library | Full + Learning Paths | Full + Learning Paths | Full + custom content |
| Sector benchmarks | 1 sector | 5 sectors | All sectors | All sectors | All + custom benchmarks |
| Admin portal | — | Basic | Full | Full | Full + SSO (OIDC) |
| API access | — | — | Read-only | Read-only | Full CRUD |
| Support | Community | Email (48h) | Chat + quarterly review | Chat + quarterly review | Dedicated CSM + SLA |
11.2 Cancellation
Cancel any time. Access remains until end of billing period. Account auto-reverts to Starter — your data is preserved, you simply lose access to paid-tier features.
12. Who this is for
12.1 Three primary personas
The AI Programme Lead at a mid-market enterprise (~500–5,000 employees). Reports to the COO or CIO. Owns the AI strategy roadmap. Needs E-ARI to:
- Demonstrate progress to executives quarter over quarter
- Translate readiness gaps into specific investment proposals
- Show the board the company isn't behind peers
The DPO / Head of Compliance at any organisation deploying AI in EU jurisdictions. Needs E-ARI to:
- Auto-classify AI systems against the Act
- Maintain a living evidence vault that survives an audit
- Generate FRIAs and Technical Files without retaining a Magic Circle law firm
The Internal Auditor preparing for an external audit, ISO 42001 certification, or a Notified Body assessment. Needs E-ARI to:
- Produce a single Submission Pack with checksums
- Show coverage trends over time (gap closure rate)
- Provide tamper-evident logs for any artefact change
12.2 Who this isn't for
We're explicit about this. E-ARI is not the right tool if you need:
- A model registry (use MLflow, Weights & Biases, or an MLOps platform)
- A general-purpose GRC tool spanning SOC2, ISO 27001, HIPAA (use Vanta, Drata, Secureframe)
- Legal advice (E-ARI surfaces obligations; it doesn't replace counsel)
- Project management of AI delivery (use Linear, Jira, Asana)
The tighter the scope, the better the tool.
13. Roadmap
Current focus areas (subject to change):
| Quarter | Theme | Highlights |
|---|---|---|
| Q2 2026 | Connector ecosystem — shipped | Evidence connectors for Slack, Jira, Confluence, Notion, GitHub, Google Drive, Microsoft 365, ServiceNow, Entra ID and Okta, with maturity stated per connector |
| Q3 2026 | ISO 42001 alignment — shipped | AI Act obligations cross-mapped to ISO/IEC 42001 and NIST AI RMF controls, in both directions |
| Q4 2026 | Multi-jurisdiction | UK AI regulatory framework, US Executive Order, NYC AEDT compliance |
| Q1 2027 | Notified Body workflows | Pre-conformity assessment templates, automated NB liaison packets |
Roadmap is published on the platform and updated monthly.
14. Glossary
AI Act — Regulation (EU) 2024/1689 of the European Parliament and of the Council laying down harmonised rules on artificial intelligence. Entered into force 1 August 2024.
Annex III — The list of high-risk AI use cases enumerated in Annex III of the AI Act (biometrics, education, employment, essential services, law enforcement, migration, justice).
Annex IV — The structure of the Technical File required for high-risk AI systems.
Article 5 — Prohibits specific AI practices: subliminal manipulation, exploitation of vulnerabilities, social scoring, real-time biometric identification in public spaces.
Article 27 — Mandates a Fundamental Rights Impact Assessment for high-risk AI systems used by public authorities or operators of essential services.
Article 50–53 — Transparency obligations for limited-risk systems (chatbots, deepfakes, biometric categorisation, emotion recognition).
Conformity Assessment — The procedure by which a high-risk AI system demonstrates compliance with AI Act requirements before being placed on the EU market.
Deployer — Any natural or legal person, public authority, or agency using an AI system under its authority. Distinct from Provider (the entity that develops or has the system developed).
DPIA — Data Protection Impact Assessment under GDPR Article 35. Often a prerequisite for FRIA.
FRIA — Fundamental Rights Impact Assessment under AI Act Article 27.
GPAI — General Purpose AI Model (foundation models). Subject to a separate obligation regime under Articles 51–55.
Notified Body — An organisation designated by an EU Member State to assess conformity of certain high-risk AI systems before they are placed on the market.
Pillar — One of E-ARI's eight readiness dimensions (Strategy, Data, Technology, Talent, Governance, Culture, Process, Security).
Pulse Check — A short monthly assessment targeting drift-sensitive questions.
Submission Pack — The single .zip export bundling Annex IV Technical File, FRIA, evidence, and manifest — ready for a Notified Body or auditor.
Technical File — The dossier required by Annex IV for high-risk AI systems. Must be kept up to date for the lifetime of the system.
Appendix A — Key URLs
| Surface | URL |
|---|---|
| Marketing site | https://www.e-ari.com |
| Pricing | https://www.e-ari.com/pricing |
| Sign in | https://www.e-ari.com/auth/login |
| Portal | https://www.e-ari.com/portal |
| Use cases (compliance) | https://www.e-ari.com/portal/use-cases |
| Evidence vault | https://www.e-ari.com/portal/evidence |
| Privacy policy | https://www.e-ari.com/privacy |
| Terms of service | https://www.e-ari.com/terms |
| Status & health | https://www.e-ari.com/api/health |
Appendix B — Contact
| Topic | |
|---|---|
| General questions | hello@e-ari.com |
| Sales / Enterprise | sales@e-ari.com |
| Privacy / DPO | privacy@e-ari.com |
| Security disclosures | security@e-ari.com |
| Support | support@e-ari.com |
E-ARI is operated by the E-ARI team. Application compute is pinned to the EU (Vercel, Frankfurt); the database runs on Supabase (AWS eu-central-2, Zurich). This document is authoritative as of the version date above; the most current version is always served at https://www.e-ari.com/docs/e-ari-handbook.md.