OSS Agent Lab · Apache-2.0

A deterministic enterprise environment for AI agent evaluation

Test whether agents act on current, resolved business context as decisions are superseded, deals move, incidents open and close, and customers write in. Meridian Works is a synthetic company with a ledger of every act, so what was true at any instant is known.

License
Apache-2.0
Runtime
Node 24+
Dependencies
None
Network calls
None

github.com/Rithmo-Inc/rithmo

Meridian Works

A synthetic company that changes over time, not a static dataset

The repository ships Meridian Works, a fictional company that sells scheduling and dispatch software to construction, facilities, trucking and field-service businesses. It has a CRM, a staff roster, a support inbox, a declared authority model, and an append-only record of every act anyone has submitted.

A seeded runner moves the company forward one business day at a time. Opportunities open, deals move and close, renewals come due, customers write in, product capabilities break and get fixed, and accounts go on a risk register. Each change is an act that passes through the same controller, is admitted or rejected with a reason, and lands on the ledger. Company state is derived from that ledger, so you can ask what it looked like at any instant since the run began.

Structural authority

Who may decide what is a predicate over a declared charter: roles, authority scopes, a 30% discount ceiling, pipeline stages, support categories and severities. The controller enforces it. It is not documentation.

Deterministic controller

Every act is admitted or rejected with a verdict and a reason: REJECTED_AUTHORITY, REJECTED_SCOPE, REJECTED_PREREQ or REJECTED_EFFECTIVE_TIME. The verdict is a function of the act, the charter and the ledger.

Append-only ledger

Every row is fsync-appended as JSONL and replayed on open. Rejected and needs-review acts stay on the record. Nothing is rewritten or deleted.

First-class supersession

A later admitted act records the act it displaces. Both stay on the ledger with the instant each took effect. The controller derives supersession from the log; an actor cannot declare it.

Synthetic business time

A seeded business-day clock and intraday planner. No weekends, no holiday model, and no wall-clock reads in the state machinery.

Reproducible seeded runs

The same seed and inputs produce the same company. Running --days=1 thirty times and --days=30 once produce the identical company.

Durable company state

Checkpointed CRM and support state, schema versions that are refused rather than guessed at, and atomic temp-and-rename writes. State lives under var/; --fresh archives it, never deletes it.

As-of reconstruction

Point-in-time views of any deal, account, support request, product incident or risk call, plus a first-response SLA read-model, at any instant since the run began.

Synthetic CRM and support inbox

223 companies, 600 contacts, 470 deals, 3,343 activities and a 32-person roster. Every address and domain uses the reserved .invalid TLD, checked by a test on every run.

Provenance on every row

Each ledger row carries the actor, the act, the verdict, the reason, the effective time, the submitted and published times, and the act it supersedes.

Every person, company, email address, domain, deal and dollar figure is synthetic. Some names are plausible because a realistic evaluation world needs them to be; none describes a real organization or person.

Why this exists

The hard failure is temporal, not retrieval

Most agent evaluation tests whether a model can find information or reason over the context it was given. That matters, but it misses a common way agents fail inside companies: the agent retrieves something that was true, and acts on it after it stopped being true. The retrieval worked. The truth moved.

  • A discount was approved, then a later decision superseded it.
  • A deal moved to a new stage after the agent read the CRM.
  • A product incident was fixed, so “that capability is broken” stopped being the right answer to give a customer.
  • A support request arrived after the agent’s snapshot was taken, and its response deadline was already running.

Each case is a question about what was operative at a specific instant, not about what some source says. That is hard to test against production systems for two reasons. You cannot replay a real company, and you rarely have a trustworthy record of what was knowable when the agent acted. Without that record, a grader cannot tell a correct action from one that happened to match yesterday’s answer.

Current, resolved business context is the answer the company has actually settled on at that moment, after any later decision has superseded an earlier one, rather than whatever a source happens to say. Acting on it is part of AI agent reliability: an agent can reason correctly and still act wrongly on a fact that has stopped being true. Showing afterwards that the agent relied on context that was operative and authorized when it acted is part of AI agent governance.

Meridian Works gives you both: a company you can replay exactly, and a ledger that determines what was true at every instant since the run began.

Known ground truth

What was true at time T is a function, not a judgement

To grade an agent’s action you need to know exactly what was operative when it acted. In Meridian Works nothing writes “the current answer” directly. Operative state is derived from the ledger:

state(T) = fold(charter, admitted acts ≤ T)

src/controller/operative.ts

charter

The declared roles, authority scopes and rules. Every fold starts from it, and every admission is checked against it.

admitted acts

Only rows with verdict ADMITTED are folded. Rejected acts, non-decisional messages and NEEDS_REVIEW rows stay visible on the ledger but never become operative truth, so an uncertain interpretation cannot become authoritative by default.

≤ T

An act counts from its effective business instant, not from when it was published. Acts that share an instant are ordered by (effectiveAt, actId), a total order, so the answer never depends on how a file happened to be read.

fold

CRM history is replayed through the same transition functions the CRM transport calls when an event happens, so there is no second implementation of history to drift from the live one. Support replies are the one exception: they do not pass through the controller, so they are folded from their own append-only attempt history, which a test reconciles against the ledger.

Because as-of views replay the real transition functions, an as-of view cannot drift from the live one by construction. A test folds the entire ledger and compares the result field by field with the company state on disk. The window is the run: the company starts from a fixed snapshot with no history of its own, so reconstruction covers any instant since the run began, not earlier.

Supersession, step by step

An illustration of the controller’s rules on one deal. The roles, verdicts and reason strings are the controller’s own; act ids are shortened. Discount decisions are a Core act, exercised directly against the controller as the test suite does. The daily runner supersedes through stage changes and incident fixes.

  1. 09:00 · A1 · account_exec

    request_discount D1 20%

    NON_DECISIONAL

    a discount request is a proposal, not a decision

  2. 09:30 · A2 · account_exec

    decide_discount D1 20%

    REJECTED_AUTHORITY

    role account_exec is not authorised to emit decide_discount

  3. 10:00 · A3 · vp_sales

    decide_discount D1 20%

    ADMITTED

    authorised, in scope, prerequisite satisfied

  4. 14:00 · A4 · vp_sales

    decide_discount D1 35%

    REJECTED_SCOPE

    decide_discount with pct=35 exceeds the scope granted to vp_sales

  5. 16:00 · A5 · vp_sales

    decide_discount D1 15%

    ADMITTED

    authorised, in scope, supersedes A3

state(09:45).D1

no discount

A1 is a proposal and A2 was rejected; neither folds

state(12:00).D1

20%

A3 is admitted and effective

state(15:00).D1

20%

A4 is on the ledger but rejected, so it never becomes operative

state(17:00).D1

15%

A5 supersedes A3; both remain on the record

src/controller/ledger.ts (excerpt)
export interface LedgerRow {
  actId: string;
  actorId: string;
  body: ActBody;
  destination: Destination;
  verdict: Verdict;
  reason: string;
  // Business effective time. Only meaningful for decision acts.
  effectiveAt: number | null;
  submittedAt: number;
  publishedAt: number | null;
  sourceAvailableAt: number | null;
  supersedes: string | null;
}
src/actions/types.ts (excerpt)
export type Verdict =
  | "ADMITTED"
  | "REJECTED_AUTHORITY"
  | "REJECTED_SCOPE"
  | "REJECTED_PREREQ"
  | "REJECTED_EFFECTIVE_TIME"
  | "NON_DECISIONAL"
  | "NEEDS_REVIEW";

// NEEDS_REVIEW is never folded into operative state. Uncertain interpretation stays
// flagged for human adjudication; it does not become authoritative truth by default.
export const AUTHORITATIVE_VERDICTS: ReadonlySet<Verdict> = new Set<Verdict>(["ADMITTED"]);
Architecture

Three layers, and the dependency arrow points one way

Each layer may import the layers below it and never the ones above. The boundary is enforced by tests/separation.test.ts on every npm test, not described and hoped for.

  1. Living Company

    May import World State and Core. Nothing imports it.

    the runner: src/living/livingRun.ts, livingStep.ts, eventFamily.ts and the families

  2. World State

    May import Core. May not import Living Company.

    src/living/worldClock.ts, dayPlan.ts, livingCrmState.ts, livingSupportState.ts, asOf.ts, supportSla.ts

  3. Core

    Imports nothing but itself and node: built-ins. Runs on its own.

    src/controller/, src/charter/, src/actions/, src/logging/, src/employees/

Living Company

What moves the company forward: five weighted event families, the consequences that follow (an incident fixed, a renewal due, a risk call, a recovery plan), the incident and risk state those own, and four local transports that publish through Core’s action client.

World State

Synthetic time, durable company state, and historical reconstruction: the seeded clock and planner, checkpointed CRM and support state, and as-of views at any instant since the run began.

Core

Authority, admission, the append-only ledger, the act vocabulary, transports, logging, the employee runtime and the model contract.

Why the boundary matters

Core runs with nothing else present, so the admission rules and the ledger can be lifted into your own harness without the runner. History in World State cannot depend on how the runner happens to generate events. And the runner cannot change company state except by publishing acts through Core’s action client, where they are admitted or rejected like any other act.

Working a customer’s support request is the one thing that needs judgement rather than a seeded draw, so it is an optional injected callback that the runner never imports.

How it is checked

  • Every import in each layer resolves inside that layer’s allowed set or to a node: built-in.
  • No file is declared in two layers.
  • Neither lower layer reaches up.
  • Every .ts file under src/ belongs to exactly one declared layer.
  • scripts/ resolves entirely inside the published tree.
  • Controls plant violations and require the scanner to fail, so a broken scanner cannot pass silently.

The same scanner fails any bare package import, which is how zero third-party dependencies is enforced. It is a pattern check over source text, not a module resolver: an import in an unusual form, or a specifier computed at runtime, is outside what it inspects.

What you can test

Evaluation scenarios the shipped code supports

Meridian Works gives you the environment and the answer key. You bring the agent and the grader. Each scenario below names the part of the repository that makes the right answer knowable.

Supersession

Does the agent recognize that an earlier decision is no longer operative?

Discount decisions, deal stage changes and incident resolutions each record the act they supersede. Both rows stay on the ledger, so you can check which one the agent relied on.

Authority and scope

Can the agent tell a valid decision from an act taken without authority or outside its scope?

The charter is enforced at admission. An out-of-scope act, such as a discount above the 30% ceiling, is on the ledger as REJECTED_SCOPE with a reason, and never becomes operative.

Temporal correctness

Given an action at instant T, what was actually operative at T?

As-of views of deals, accounts, support requests, incidents and risk calls at minute resolution. Events that share an instant have a stable, total order.

Stale context

Does the agent keep acting on a fact after the company has moved on?

Give the agent a view from T0, advance the company, and compare what it does at T1 against state(T1). The advance is seeded, so the change you test against is the same on every run.

Ownership-bound acts

Does the agent route work to the person who actually owns it?

A recovery plan must come from the owner of the risk call it answers, and renewal pricing from the owner of the renewal. Anything else is REJECTED_PREREQ. Owners are assigned and enforced; the current runner does not reassign them mid-run.

Support handling against an SLA

Does the agent answer a customer correctly, and on time, given what was true when it answered?

runLivingWorld takes an optional processSupport callback, the one place the company needs judgement. The SLA read-model reports MET, MISSED or UNANSWERED as of any instant.

Reproducibility

Can the same evaluation scenario be replayed exactly?

Identical inputs produce an identical ledger, delivery record and act ids. The test suite checks this, and checks that the controlled clock and injected id source are load-bearing.

What is not in the box

There is no bundled agent, scoring harness or benchmark leaderboard. The CLI connects no agent, so the support requests it raises are recorded and stay open rather than being answered by a stub. The staff, products and documents are fixed today; what evolves is the pipeline, the support inbox, incidents, renewals and the risk register.

Run locally

Clone it and run the company forward

Requires Node.js 24 or newer. Node runs the TypeScript sources directly, so there is no install step and no build step. package.json declares no dependencies and there is no lockfile; node --test is the whole toolchain.

shell
$ git clone https://github.com/Rithmo-Inc/rithmo.git
$ cd rithmo
$ npm test
$ npm run company -- --days=30

npm test runs the full suite. On a fresh clone two reconciliation tests report skipped: they compare a live company against its own ledger, and there is no company yet. After a run they pass.

npm run company writes state under var/, and the next invocation continues from there. --fresh moves the existing company into var/archive/ and starts a new one. --seed=<n> and --start=<YYYY-MM-DD> apply when a company is initialized; the defaults are seed 1 and 2026-10-01.

npm run company -- --days=30 on a fresh clone, default seed (excerpt, verbatim)
=== MERIDIAN WORKS ===
  initialising  at 2026-10-01, base seed 1
  running       30 business day(s)
  support       no agent connected — raised requests are recorded and stay open

=== 30 BUSINESS DAY(S) ===

  2026-10-01 Thu  day 1  —  1 event(s)  [tempo 0.3, planned 1]
    11:43  MOVE      MW-D-0441  qualifiedtobuy -> presentationscheduled
              Fairmont Freight - Renewal  $50,124  owner MW-EMP-06  act MW-ACT-1-d0001-001  ADMITTED
    17:59  RENEWAL   MW-LD-0001 opened on MW-CO-0020, due for renewal
              Rivermark Field Services - Renewal  $72,284  opens at appointmentscheduled  owner MW-EMP-05  act MW-ACT-1-renew-MW-CO-0020-2026-10-01  ADMITTED
  …

=== WHERE MERIDIAN IS NOW ===
  synthetic date 2026-10-01  ->  2026-11-11
  business day   0  ->  30
  base seed      1
  opportunities opened 26
  deals advanced       114
  closed won           15
  closed lost          15
  …
  support requests     57
    caused by incident 3
    independent        54
    open, unanswered   57  (no agent is connected; nothing was answered and nothing was faked)
  product incidents    2 started, 2 fixed

Same seed, same output. Run it again with --fresh and the company is identical.

Status: version 0.1.0, early. The surface above is tested and stable enough to build evaluations on, and all three layers will grow. Determinism and the layer boundary are the two properties intended to hold across changes. One known limitation: the Customer Success risk trigger is not yet calibrated, so a 30-day run usually leaves the risk register empty, and the runner says so.

Scope

The evaluation substrate is included. Inferring truth from real evidence is not.

This repository is the end of the problem where the right answer is already known. Truth here is declared and controlled, which is what makes it useful for evaluation, and also why it leaves out the hard parts of working inside a real company.

Included

Declared organizational truth
Roles, authority scopes, a discount ceiling, pipeline stages, support categories and severities, as predicates the controller enforces.
Deterministic controller and ledger
Every act admitted or rejected with a verdict and a reason, appended to an immutable log. Replaying the log reproduces the state exactly.
Synthetic time
A seeded business-day clock and intraday planner.
Durable company state
Checkpointed CRM and support state with versioned schemas and atomic writes.
Historical reconstruction
As-of views of any deal, account or support request since the run began, and a first-response SLA read-model.
A synthetic CRM dataset
Companies, contacts, deals, activities and a staff roster in seed/hubspot-manifest.json.
A deterministic runner
npm run company advances Meridian one business day at a time. Seeded, offline and reproducible.
A model contract, with no provider
Request and reply shapes, a ModelClient interface and a scripted client for tests. No HTTP client, no API-key read, no network call.

Deliberately not included

Inference from real evidence
Meetings, transcripts, threads, documents, tickets and the systems they live in. Here, truth is declared and controlled.
Uncertainty resolution
Deciding what a decision was when the evidence is partial, contradictory or never stated outright. Here, an act is structured and its verdict is a function.
Production integrations
No connector, no OAuth, no provider client, no live transport. Only the authorization gates a live transport would have to pass.
Governance and operations
Tenancy, access control, retention, audit and administration.

What generalizes from the repository is the evaluation substrate: deterministic state, structural authority, supersession, and as-of reconstruction.

Commercial Rithmo

Testing known truth is one problem. Determining truth inside a real company is another.

In Meridian Works the right answer is known before the agent acts, because truth is declared and every act is structured. A real company has no ledger like that. Decisions are made in meetings, threads and calls, and the systems agents read from often still hold the previous answer.

The commercial Rithmo product is the fact-checker for AI agents. It works from a company’s real evidence to establish whether the business context behind an agent’s action can still be trusted, and lets agents check before they act. This repository is not that product and does not contain it.

Build an evaluation on known truth

The repository is public under Apache-2.0. Issues and questions are welcome.

Or email the builders at hello@rithmo.ai.