Heuristic Research Systems

probo“I prove”

An agentic research methodology that can show where every conclusion came from.

Putting a research question to a language model produces an answer with no provenance: fluent, plausible, and impossible to audit. probo is the opposite arrangement: coordinated agents working a defined brief, evidence kept separated by origin, scored against versioned schemas, and released only by a person. The methodology is the product. What it produces in any given market is an instance of it.

What makes it a method rather than a prompt

Four properties, each of which exists to make a finding contestable. A conclusion nobody can challenge is not a research output.

  1. Agents gather

    Coordinated, not prompted

    A swarm of agents works a defined research brief, each routed to the model suited to its task. This is not one model answering one question. It is a division of labour, scheduled and repeatable, over a corpus no analyst could read in the time available.

  2. Sources stay separate

    Evidence is kept apart by origin

    What a subject says about itself and what independent sources say about it are gathered and scored as different things. Merging them early is how bias becomes invisible, once combined, no reader can tell which claim came from where.

  3. Schemas are versioned

    A score means something specific

    Every assessment records the schema version it was made against. Definitions can improve without silently rewriting history: a changed schema produces a new pass, not a quiet edit to the old one.

  4. Cadence, not campaigns

    Research that does not go stale

    Assessment runs on a schedule rather than as a one-off engagement, so findings track a moving market instead of describing the week they were commissioned.

The disagreement is the finding

The method scores the same criteria more than once, from classes of source that are kept independent of each other, and then publishes the passes side by side rather than reconciling them into a single number.

Where the passes agree, confidence is high and the reader can move on. Where they diverge, something is worth knowing: a claim not carried by the independent record, or strength the subject has not thought to assert. A blended score erases exactly that signal, which is why it is the industry default and why probo does not do it.

The cost is honest: separated passes are harder to read than a ranking. The gain is that a finding survives being questioned.

Where the two passes disagree

An illustration, not an assessment. What a subject publishes about itself and what independent sources record are gathered and scored separately, and never see each other. Plotting them together is the first time they meet.

0100
  • Data handling42 apart
  • Access control14 apart
  • Incident response7 apart
  • Third-party oversight48 apart
  • Retention and deletion23 apart
What the subject publishes about itselfWhat independent sources record

The bar between them is the finding. Convergence is a confidence signal. Divergence is the result worth reporting, and it only survives because the passes were never merged.

What an average would leave you with

One dot per dimension, every gap above gone, and no way to recover which reading came from where. Easier to read, and strictly less informative than either input.

Illustrative divergence between self-published claims and independently observed evidence
DimensionClaimedIndependently observedGap
Data handling925042
Access control617514
Incident response74817
Third-party oversight884048
Retention and deletion704723

Three kinds of research, one method

The four properties hold whatever is being investigated, but they do not carry equal weight. Each kind of work leans hardest on a different one, and knowing which is most of the craft.

Market research

Who can actually do this, and how would we know?

Assessing capability across a field of providers or a category. The separation that matters here is between what a subject publishes about itself and what independent sources record, because the two are the most likely to disagree and the disagreement is the finding.

Leans hardest on

Cadence

Markets move. An assessment that describes the quarter it was commissioned in is a historical document, so the work runs on a schedule rather than as an engagement.

Academic research

What does the evidence actually support?

Synthesis across a literature, where the sources are papers rather than claims. Separation by origin becomes the distinction between primary evidence, secondary interpretation, and assertions that have been cited so often they are mistaken for findings.

Leans hardest on

Reproducibility

A synthesis nobody else can re-run is an opinion with a bibliography. Recording the schema and the sources behind each judgement is what lets a second researcher reach the same place or show where it breaks.

Advisory frameworks

What should other people use to decide?

Producing the model itself: the capability structure, maturity scale or reference architecture that others will apply to their own situation. Here the schema is not an instrument of the research, it is the output.

Leans hardest on

Versioning

A framework that changes quietly invalidates every assessment ever made with it, and nobody can tell which. Versioning means a framework can improve without rewriting its own history.

Confidence is cheap. Provenance is not.

Any system can produce a fluent answer. The question is whether it can show you what the answer rests on.

Agentic, not autonomous

Agents do the gathering and the scoring. They do not decide what the world is told.

Human-gated release

Collection is continuous; publication is not. No finding reaches a reader without a person releasing it. Automation that publishes itself is how unreviewed claims enter the record.

Traceable to source

Each score carries the schema version and the class of evidence behind it, so any conclusion can be walked back to what produced it.

Reproducible

A run can be repeated against the same schema and compared, which makes a disputed finding testable rather than a matter of authority.

For questions where being wrong is expensive

The method is domain-agnostic. What it needs is a question worth defining precisely and a decision that someone will later be asked to justify.

Why we build

A finding is only worth as much as the account it can give of itself. Ours are built to be argued with: to name what was looked at, where it came from, and what would have changed the answer.

Independent of source, versioned in method, released by a person.