Product Decision System of Record · by Helium

Product decisions need more than an agent.

pm.gold is the governed system of record for product evidence, decisions, and outcomes. Use ChatGPT, Codex, Claude, or any model to do the work — pm.gold preserves what supported it, who approved it, what shipped, and whether it worked.

Supported claims cited — gaps shown, not inventedAgents propose · a human approves · the record is append-onlyModel-agnostic — your record outlives any one vendor
app.pm.gold
The pm.gold workspace: sources traced, the commit gate, and quick actions — a real screenshot of the product.
Entra & Google SSOImmutable auditTenant isolation, fail-closedPermission-faithful retrieval

What only a system of record can do

Four questions an agent can’t answer for you.

Not because the model isn’t good enough — because the answer isn’t in a conversation. It’s in a governed record.

Defend any decision

“Why did we build this, and who approved it?”

Walk any requirement backwards: the evidence it cites, the sections that honestly had none, the precedent that informed it, every version, and the human who committed it — with separation of duties on the record.

Intent vs. reality

“Did we ship what we approved?”

Compare the approved PRD clause by clause against implementation evidence. Each gap is filed as a proposal at the gate — surfaced, not quietly lost between the spec and the sprint.

Learn from decisions

“Have we tried this before, and did it work?”

A new draft recalls similar past decisions and their recorded outcomes — won or lost — and the corrections your team keeps making. Precedent is linked to the draft it shaped, so the learning is auditable.

Catch the drift

“What did we decide that is no longer true?”

Committed decisions keep watching the evidence beneath them. When a cited source moves, every decision still standing on the old version is listed — by query over the record, not by a model’s judgment. Each arrives as a proposal at the gate, capped by a weekly interrupt budget so it stays signal.

Each answer is assembled from committed rows — never generated. Where a step hasn’t happened yet, pm.gold says so rather than implying it.

The problem

Your agents are getting better. Your decision record isn’t.

ChatGPT and Claude now remember context and cite connected sources. What none of them keep is a governed record of what your team actually approved — the evidence behind it, who signed off, what shipped, and whether it worked. As agents write more of the work, that record is what gets scarce.

48%

of a PM’s time goes to unplanned firefighting.

Product Focus, 2024

23%

use AI where they need it most — shifting priorities.

Atlassian, 2026

84%

worry their product won’t succeed in the market.

Atlassian, 2026

How it works

One loop — and it doesn’t end at the record.

Agents do the work; the Brain holds the truth. Every step is typed, versioned, and traceable to the source that drove it — and the last step points back at the third.

01

Capture

Meetings, calls, docs, tickets — ingested as evidence.

02

Reason

Agents cluster signals and draft, grounded in the Brain.

03

Propose

PRDs, scores, decisions — supported claims cited, gaps flagged.

04

Gate

A human approves or rejects. Nothing is automatic.

05

Record

Approved work enters the append-only Brain.

06

Watch

Decisions keep watching their evidence. Change returns to step 03.

Nothing enters the record until a human approves it. The commit gate is the product, not a setting.

The record checks itself

Evidence changes. Most records never notice.

A decision is only as good as what it rests on. Six months later the pricing page it cited has changed and the competitor it assumed has moved — and the document still reads exactly as confident as the day it was approved. That isn’t a documentation problem. It’s the failure mode of every decision record ever kept.

pm.gold commits a decision with its assumptions attached. Each one names what would have to be true, and what would prove it false. Watched pages are re-fetched on a schedule and diffed. A change that matters becomes a proposal at the same gate everything else goes through — nothing is edited automatically, and nothing is quietly dropped.

It’s allowed to interrupt you a fixed number of times a week. That budget is the feature, not a limitation of it. A system that can raise an alarm whenever it likes becomes noise inside a month — and noise is worse than silence, because silence is at least honest about how much it knows.

The Contradiction Ledger

A decision cited version 1; that source is now on version 4. Finding the decisions standing on the old evidence is a query over the record — not a model’s opinion about whether something still holds.

Ledger Health

What share of your committed decisions still rest on current evidence — and the negative space beside it: the objectives you hold no evidence about at all. A record can be wrong by being silent.

It learns what you ignore

What you dismiss at the gate feeds back into what gets surfaced. The calibration can only ever damp a category, never amplify one — so the loop can quiet itself down and has no path to shout.

Once a week, the same record tells you what’s waiting for you and what has gone stale — to your own inbox, or your own Slack. Counts, titles and links; never the evidence text.

Capabilities

Everything a PM touches — captured & cited.

See all capabilities →

Meeting capture

Transcripts turned into decisions, actions, and PRD-ready snippets — each timestamped to its source.

PRD generation

Draft from evidence with inline citations, then red-team and ship-check before it ships.

The commit gate

Agents propose; a human approves. Every decision and override recorded in an immutable audit.

Why pm.gold

Not another tool. The most trusted voice in the room.

A typed decision graph

Not notes in a chat history: meetings, signals, decisions, PRDs, risks and outcomes as typed nodes with append-only versions and traversable evidence chains. Every artifact knows what it rests on.

Learns from your approvals

pm.gold captures what humans change between the draft and the approved version, distills governed calibration rules, and applies them to future work. Early controlled tests reduced later editing while grounding held flat — directional, and we report it that way.

Your record outlives the model

Models swap behind a clean seam, so every leap in AI is an input rather than a migration — and your decision history never lives inside one vendor’s memory.

Security & trust

Enterprise-grade by construction, not by checklist.

The guarantees are enforced in the data layer — so they hold even when something above them has a bug.

SSO — Entra & Google

Sign in through your identity provider, bound to your tenant.

Immutable audit

Every decision and override is recorded and can’t be edited or deleted.

Tenant isolation, fail-closed

Enforced in the data layer — row-level security on every content table, failing closed by default. Dedicated single-tenant deployment available for Enterprise.

Grounded & defensible

Every AI claim cites its source — nothing asserted that can’t be traced.

FAQ

Questions, answered honestly.

The short version. For the full, honestly-labeled security posture, see the Security page.

What is pm.gold?
It’s the governed system of record for product evidence, decisions, and outcomes. Your meetings, docs, and tickets become cited evidence; agents propose PRDs, scores, and decisions from it; a human approves at a commit gate; and what’s approved becomes a versioned, traceable record — with the evidence, the approver, what shipped, and whether it worked all preserved.
How is it different from ChatGPT, Codex, or Claude?
They’re powerful agents for researching, drafting, and executing work — and they now remember context and cite connected sources too. pm.gold is the durable decision layer underneath them: a typed, permission-faithful record of evidence, approvals, versions, and outcomes. Your agents can propose work into pm.gold, but only what a human approves becomes organizational knowledge. Use whichever model you like — the record outlives the choice.
Does it make things up? How do I trust the output?
Every supported claim carries a citation to the source it came from. Where the evidence doesn’t support a section, pm.gold shows the gap instead of inventing support — citations can only ever point at retrieved sources, so a fabricated citation isn’t expressible. We publish an eval harness rather than claim perfection. And nothing enters the record without a human approving it at the commit gate.
Who is it for?
Product managers and product teams who need their signals, decisions, and rationale to stay connected, recallable, and defensible instead of scattered across a dozen tools.
How do we get our context in?
Attach documents directly — PDF, Word, PowerPoint, Excel/CSV, email, or a .zip of them — paste notes, or capture a public web page. All of it lands as cited, groundable evidence, and company-wide context stays separate from personal context. Connectors to systems of record are labeled honestly by state: adapters (permission mapping built, scheduled sync still staged), and roadmap. We don’t call a connector live until its sync is proven against a real tenant.
Is our data used to train AI models?
No. Your content is not used to train generalized models — enforced by the no-training terms of the commercial model APIs we use, and committed contractually in our DPA. Models are reached behind a model-agnostic seam, and each call gets only the data the task requires.
Can we use pm.gold from Claude or ChatGPT?
Yes. pm.gold is an MCP server, so you can add it to Claude, ChatGPT, Codex or any MCP client and query the record from where your team already works. An admin issues a key, and the agent gets four read tools — find precedent, read an approved decision, fact-check a document, compare intent to what shipped — plus three that can only propose, including one for research: your assistant searches with its own tools and hands us a URL, we fetch the page ourselves, and it enters as evidence awaiting review. The key is bound to the person who issued it and inherits exactly their permissions, so an agent can never retrieve a source its owner cannot see. And there is no approve tool for it to call: nothing an agent drafts enters the record until a human commits it.
Will it tell me when a decision we made is no longer safe to rely on?
That is what the sixth step of the loop is for. A committed decision carries its assumptions, each naming what would prove it false; the pages those point at are re-fetched on a schedule and diffed. When a cited source changes, the decisions still standing on the old version are found by query — not by asking a model whether something still holds — and each one arrives as a proposal in your review queue with the diff attached. Nothing is rewritten automatically. And there is a cap on how often it may interrupt you in a week, because an alerting system without a budget becomes a filter rule inside a month.
How is access controlled? Is our data isolated from other companies?
Tenant isolation is enforced in the data layer and fails closed by default — a request without verified workspace context returns nothing. Role-based access separates what you can do from what you can see, and source-level permissions are enforced at retrieval. A dedicated single-tenant deployment is available for Enterprise.
How do we sign in?
Single sign-on through Microsoft Entra ID or Google, bound to your tenant — we never see passwords. Magic-link and password options are also supported.
Which AI models does it use?
pm.gold is model-agnostic — models swap behind a clean seam, so you’re always on the frontier with no lock-in to one provider. Bring-your-own-key is supported for teams that prefer their own provider relationship.
Are you SOC 2 certified?
Not yet, and we say so plainly. SOC 2 and an independent penetration test are planned, not complete. A DPA is available today, we complete a CAIQ-Lite questionnaire on request, and our subprocessor list is published — see the Security page for the full, honestly-labeled status.
Can it connect to our existing tools?
You connect the way you already do in Claude or ChatGPT — authorize access in the browser; we never take a password. Each connector carries every item’s real permissions into the Brain, fail-closed. We label them in four honest states: live (sync proven against a real tenant),design partner, adapter (permission mapping built, scheduled sync still staged), and roadmap. Today the built connectors are adapters — we don’t call one live until it syncs your tenant.
How do we get started?
We’re onboarding design-partner teams now. Request a demo and we’ll stand up your workspace.

Still have questions? Request a demo →

The gold standard

Give your best decisions a gold standard.

We’re onboarding design-partner teams now. Book a walkthrough and we’ll stand up your workspace.