Archive-to-Royalties: Paying the People Who Filled the Archive

When a conference talk answers questions for a decade instead of an afternoon, the economics of who pays whom have to change. The infrastructure exists. The choice is yours.
TL;DR. Conference archives are becoming machine-queryable, so use becomes unbounded while contributor compensation stays at zero. We call the fix archive-to-royalties: metering each use of a contributor's material, attributing it, and paying on published terms. There are three paths — credit, rent, and ownership — and the contributor elects, not the platform. Attribution is not a solved problem. That is the honest caveat, and it is why the published formula matters more than the optimal one.
A spine surgeon spends thirty hours preparing a talk on revision strategy for proximal junctional failure. She delivers it once to a room of two hundred colleagues at a national congress. The session is recorded. The recording goes into the society's on-demand archive, behind a subscription that costs members €450 a year. The surgeon's compensation is a flight, a hotel, and a line on her academic CV.
That arrangement has been in place for decades, and nobody has seriously questioned it. It was defensible when a recording could only be watched, because watching scales linearly with human attention, and human attention is scarce. It stops being defensible the moment the archive becomes machine-queryable. A talk that is retrievable answers questions continuously, in every timezone, for a decade. Its use becomes unbounded while its compensation stays fixed at zero.
The arrangement
Flat-fee archive licensing made sense for a long time, and it is worth understanding why before arguing that it has broken.
A learned society runs a conference. Speakers donate their preparation and their time. The society records the talks and sells access to the archive as an annual subscription. The revenue funds the next conference, the society's overhead, and the educational mission. Speakers are paid in prestige — the invitation itself is a credential, and the recording extends its reach.
This model worked because every use of the archive required a human to sit down and watch a video. The number of views was bounded by the number of members, their available time, and their interest in a specific topic. A brilliant talk on C1-C2 fixation techniques might be watched five hundred times over three years — a lot for a niche subject, but still a finite, human-scale number. The society could budget accordingly. The speaker could expect recognition but not rent.
The model was also fair in a rough way. The society bore the cost of recording, hosting, and distribution. The speaker bore the cost of preparation. Both contributed something the other could not easily replace. The subscription revenue reflected aggregate human interest, and aggregate human interest was the only kind of demand there was.
What changed is not that this model became unfair. What changed is that a new kind of demand appeared that the model was never designed to measure.
Archive-to-Royalties
We need a name for what comes next, because a concept without a name cannot be cited, argued with, or adopted by someone else.
Archive-to-royalties is the conversion of a static recorded archive into a metered one, in which each use of a contributor's material is measured, attributed, and compensated on published terms. It has four requirements, and a system missing any one of them is not doing it.
Provenance at ingestion. Every retrievable unit carries a resolvable record of who made it and under which permission. No rights record, no ingestion. This is the requirement that separates the model from every scraped corpus.
Attribution at query time. The system records which material contributed to which answer, at a resolution fine enough to divide money by.
Compensation on a published formula. The split is stated in advance, applies to everyone equally, and changes only through a governance process rather than by operator decision.
An auditable record. The contributor can verify what happened without asking the platform to confirm its own arithmetic.
The term covers the transition, not just the end state. An archive that already meters is not doing archive-to-royalties; it has done it. The interesting and difficult part is the conversion of material recorded under one set of expectations into a system operating under another.

What changed
Four things arrived at once — none of them existed in production form eighteen months ago.
The first is retrieval systems that can identify which source material actually produced an answer. When a language model is grounded on a corpus of conference recordings, it does not simply retrieve a video for a user to watch. It reads the transcripts, chunks them into passages, and uses those passages to construct a novel answer to a specific question. The system can report which passages it drew on. This is not a speculative capability — it is how retrieval-augmented generation already works, and the attribution layer is improving fast.
The second is a payment protocol that lets a machine pay for a single query without an account, an invoice, or a human in the loop. The HTTP 402 status code was reserved for payments in 1997 and never used. In May 2025, Coinbase released x402, an open protocol that revives it: a server responds to a request with 402 and a payment specification; the client signs a stablecoin payment and retries the request with proof attached. The server verifies, settles, and returns the content. By April 2026, the protocol had been contributed to a neutral foundation. Cloudflare and AWS both shipped edge-level support. Visa and Stripe connected it to conventional card rails through their respective agent-commerce protocols. The infrastructure is real.
The third is a settlement layer where the record of use cannot be quietly rewritten by the party doing the paying. On-chain settlement provides a tamper-proof ledger of which query accessed whose material, when, and at what price. This matters because the history of content licensing is largely a history of opaque accounting. A settlement record that neither party can unilaterally edit is the prerequisite for any trust relationship that scales beyond handshake deals.
The fourth is the collapse of inference cost. When computation approaches free, the marginal cost of answering a question becomes dominated by what is owed to whoever supplied the knowledge. The royalty stops being a rounding error on a compute bill and becomes the main component of the price. That inversion is the strongest available argument that archive-to-royalties gets more viable over time, not less.
The retrospective problem
The harder question is not future terms but the material already recorded — under releases that predate machine use and never contemplated it. That is most of the value and all of the difficulty.
Consent obtained for one purpose does not extend to another. The honest route is to go back and ask. The reason nobody does is operational cost, not principle. A society with fifteen hundred recorded talks spanning two decades cannot renegotiate fifteen hundred individual agreements before switching on a retrieval system. This is where archive-to-royalties earns its name: the conversion is the hard part, and it requires a framework that can handle material whose original terms are silent on machine use. The practical path is a default election — credit, with an opt-up to rent or ownership — that takes effect on ingestion and can be revisited by the contributor at any time. It is not perfect. It is better than the alternative, which is to pretend the old terms were fine and proceed without asking.
Contributors who cannot be found
Some rights-holders are untraceable. Some have died — which in surgical teaching includes a disproportionate share of the most valuable material, because the canonical technique talks were given by the generation that trained the current one.
Our position: orphan material is ingestible under attribution-only, with the royalty share accruing to escrow indefinitely against a public claim register, never converting to platform revenue. Attribution survives regardless, because a name is not a payment and should never be conditional on one. A system that drops the contributor's name when it cannot find their wallet has confused two different obligations.
Three paths
These are the three payment structures available within an archive-to-royalties system. Every conference organiser sitting on a video archive now faces a choice among them — and the choice should be the contributor's, not the platform's.
Path one — Credit
No money changes hands. The contributor grants use in exchange for verifiable, machine-readable attribution. Every answer that draws on their material emits a durable record. The contributor accumulates an auditable account of where their teaching travelled — which questions, which jurisdictions, how often.
This suits a large fraction of senior academic medicine. Many contributors are at institutions that prohibit outside payments, or in countries where receiving foreign micro-payments is a tax and compliance burden out of proportion to the sums. Some genuinely prefer it. Survey work in open science and data sharing consistently finds that researchers value citations and formal recognition more than direct financial compensation — the citation is the career currency, and a machine-readable citation that tracks use across a decade is a more durable form of recognition than a footnote in a proceedings volume.
The failure mode must be stated plainly: it is free labour with a receipt. If the platform earns revenue and the contributor does not, attribution becomes a fig leaf for exploitation. The mitigation is that credit must be the contributor's election, never the platform's default.
Path two — Rent
A per-use royalty. Each query carries a price. The price splits between infrastructure, the convening body, and the contributors whose material was actually used, on a published formula. Payment is continuous, proportional to use.
This suits contributors whose material is durably useful rather than momentarily fashionable. A talk on complication management for posterior cervical fixation will answer questions for a decade. A keynote on last year's controversy will not. The long tail of the archive benefits more than the keynote — and that is the right outcome, because the long tail is where most of the teaching value lives.
The analogy is performing-rights organisations — GEMA, PRS, ASCAP, BMI. A meter runs, plays are logged, money distributes on a published formula, and the composer need not trust the broadcaster because neither of them does the counting. Raptive's strategy lead has publicly described pay-per-value as the most complex but best model and explicitly compared it to ASCAP and BMI royalties. That comparison is not ours; it is already the industry's.
Three failure modes. Concentration: a handful of topics absorb most queries and most of the pool. Variance: most contributors earn amounts too small to matter — precisely what happened to musicians under streaming. And one practical limit: a transfer of eleven euros costs more to process than it is worth, so payments accumulate to a threshold or an annual statement. Contributors should hear this from us before they discover it. Contestability: the formula decides who gets paid, so whoever controls the formula controls the outcome. It must be published, open to inspection, and changeable only through governance, not by operator decision.
Path three — Ownership
The contributor takes a stake in the archive as a whole rather than a payment for their own retrievals. Reward tracks the aggregate value of the commons, not the individual's query count.
This suits early contributors, who take the most risk. Their material bootstraps a corpus that has little value until it reaches critical mass. It is the only path that rewards the people who contribute before the archive is worth anything. It also solves an incentive problem the rent model creates: under per-use royalties, a contributor is rationally indifferent to whether anyone else contributes. Under an ownership stake, every contributor wants the archive to grow.
The failure mode: reward decouples from actual usefulness. A contributor whose material is never retrieved is rewarded identically to one whose material answers a thousand questions. Valuation is opaque. Any instrument of this kind carries regulatory weight in most jurisdictions. And it asks contributors to accept risk in place of cash, which is a real cost.
The analogue is older and duller than the technology: producer cooperatives and mutual societies. The dairy co-op, the mutual insurer. Structures that a sceptical professor finds reassuring precisely because they predate the last decade of fintech by a century.
The point
These are not competitors. They are elections, and a serious platform offers all three. The failure mode of every system built so far is that the platform picked one path on everyone's behalf, usually the one cheapest for the platform.
It is worth naming the fourth option, because it is the incumbent. The lump-sum buyout: an archive is licensed wholesale for a fixed fee, and the individuals who made it receive nothing and learn nothing. When lump-sum licensing is the only mechanism available, the price of past use ends up being set retrospectively — through negotiation under threat, or through litigation. A metered model prices use as it happens, prospectively and by agreement. That is a structural argument, and it does not require naming anyone.
The hard part
Attribution resolution is the binding constraint on archive-to-royalties, because you cannot divide money more finely than you can measure contribution. This is the section where technical readers decide whether to trust the piece.
When a system answers a question by drawing on eight sources from five contributors, how much credit does each receive? There are four approaches in the current literature, and none of them is right.
Uniform across retrieved sources. Trivial to compute, trivial to explain, and wrong. It rewards being retrieved rather than being useful.
Retrieval-similarity weighted. Credit proportional to how closely each chunk matched the query. Cheap and defensible, but similarity to the question is not the same as contribution to the answer.
Citation-grounded. Credit only where the generated answer demonstrably drew on a source. Closest to intuition; hardest to compute reliably; sensitive to how grounding is verified.
Cooperative game-theoretic. Shapley-style attribution measures each source's marginal contribution across all possible subsets of the retrieved set. Principled and fair by construction, but exact computation is exponential in the number of documents — each evaluation is another model call, so it is impractical per query. Active research on approximations is promising. Nematov et al. ("Source Attribution in Retrieval-Augmented Generation", arXiv:2507.04480, July 2025) systematically compare six approximation methods in the RAG setting and find that Kernel SHAP and ContextCite achieve above 0.95 Pearson correlation with exact Shapley using approximately 100 samples, while capturing redundancy, complementarity, and synergy between documents — the inter-document relationships that naive heuristics miss. MaxShapley (arXiv:2512.05958, December 2025) takes a different approach: by exploiting a decomposable max-sum utility function, it computes attribution in linear time rather than exponential, reaching quality comparable to exact Shapley with up to an eightfold reduction in tokens over prior methods. Crucially, the MaxShapley paper is explicitly motivated by incentive-compatible generative search with fair context attribution — the need to compensate content providers for their contributions to generated answers. That motivation is coming from the computer science literature itself, not from us.
The practical position: a production system today will use a cheap approximation as the base, with citation grounding as a correction. The honest framing is that this is a first iteration open to challenge, not a solved allocation. The published formula and the right to contest it matter more than the formula being optimal.
One further problem: redundancy. If two contributors taught the same technique in similar terms, neither is individually necessary, and a strict marginal-contribution method may credit both near zero. Any real system needs a rule for this. We do not have a good one yet.
Prior art
Components of archive-to-royalties exist in media and publishing. The unbuilt part is its application to scientific live events.
In media, content licensing for AI use has moved from one-off deals to structured marketplaces. Microsoft's Publisher Content Marketplace, announced 3 February 2026, was co-designed with publishers including the Associated Press, Condé Nast, Hearst, Business Insider, Vox Media, USA TODAY, and People Inc. Publishers set their own terms and pricing; payment is tied to delivered value rather than traffic or impressions. Yahoo was the first demand partner. Microsoft's stated rationale: individual publisher-by-publisher negotiation created too much friction and did not scale. That is our argument arriving from the opposite direction — the largest player in the space has concluded that bilateral licensing does not scale and that infrastructure is required. Amazon has been reported to be developing a comparable marketplace, alongside separate direct licensing agreements; no product has launched.
The industry now recognises four models: lump sum, pay-per-crawl, pay-per-query (pay-per-inference), and pay-per-value. Really Simple Licensing, whose 1.0 specification was finalised in December 2025, lets a publisher declare terms including pay-per-inference in machine-readable form. Reddit, Yahoo, Medium, O'Reilly, and Ziff Davis were early backers. Honest caveat: as of early 2026, no major model developer had publicly committed to honouring it.
In payments, x402 revives the HTTP 402 status code. Released by Coinbase in May 2025, contributed to a foundation in April 2026. Cloudflare and AWS shipped edge-level support; AWS added managed agent payments to its Bedrock platform; Visa and Stripe connected the protocol to conventional card rails. Cloudflare's own reporting found that by June 2026, fifty-two per cent of crawler requests were for AI training, roughly doubled from spring 2025 — the empirical basis for the claim that ads and subscriptions are failing as the traffic mix shifts from humans to agents.
The counter-evidence: independent analysis in March 2026 found daily settled volume still modest and substantially composed of testing rather than commerce. Transaction-size data shows weight shifting toward larger payments rather than true micropayments. The infrastructure is real; the demand is not yet proven.
In attribution research, the literature cited above is active and published. What does not exist: no system applies contributor-level, per-query attribution royalties to a scientific live-event archive. Medical conference archives are sold today as flat-fee on-demand subscriptions — universally. Speakers receive no continuing compensation and no usage visibility. Decentralised-science efforts have concentrated on research funding, IP ownership, and publishing — not on the education and live-event layer.
The defensible claim is narrow and true: the mechanism is proven in media, the payment rail is proven in infrastructure, the attribution mathematics is published, and nobody has assembled them for the conference archive.
What we are doing
We are in active discussion with a major spine education organisation about applying this framework to their archive. We have not signed a deal.
What is built: a retrieval system that indexes conference transcripts at chunk granularity, attaches provenance metadata to each chunk, and can report which sources contributed to a given answer. A payment integration that handles the 402 request-pay-retry handshake. A settlement record that logs each transaction immutably.
What is not built: a production attribution formula. The system currently uses retrieval-similarity weighting as a base with citation grounding as a correction. This is a first iteration. A governance process for changing the formula is designed but not instantiated. Per-token pricing — where cost reflects the actual token count drawn from a contributor's material rather than a flat per-request fee — is more honest than per-request pricing and is the model we are building toward, but it is not yet the default.
The batching threshold — the point at which individual per-query settlements are aggregated into a single on-chain transaction to keep gas costs below the payment value — is set by transaction cost on the settlement network. Below the threshold, the system holds a verifiable commitment to pay. Above it, it settles. This is a practical compromise, and the threshold will move as settlement costs move.
Close
The surgeon who prepared the talk on revision strategy for proximal junctional failure will not retire on royalties from her recording. That is not the point. The point is that her talk, and thousands like it, are becoming machine-queryable whether or not anyone decides how to handle it. The terms will be set by whoever builds the system. The choice available to a learned society is not whether its archive gets metered but whether it does the metering itself, on published terms, with its contributors electing how they are paid.
That conversion is what we mean by archive-to-royalties, and it is available now rather than eventually.
If you hold an archive and this is your problem too, we would like to hear from you.
Sources
Shapley attribution in RAG: Nematov et al., "Source Attribution in Retrieval-Augmented Generation", arXiv:2507.04480, July 2025. Compares exact Shapley with six approximation methods for document-level RAG attribution; finds Kernel SHAP and ContextCite achieve >0.95 Pearson correlation with ~100 samples.
MaxShapley: arXiv:2512.05958, December 2025. Decomposable max-sum utility gives linear-time attribution; explicitly motivated by incentive-compatible compensation for content providers in generative search.
Microsoft Publisher Content Marketplace: Announced 3 February 2026. Co-designed with AP, Condé Nast, Hearst, Business Insider, Vox Media, USA TODAY, People Inc. Yahoo as first demand partner. Payment tied to delivered value with usage reporting.
Amazon: Reported by TechCrunch (February 2026) to be developing a comparable marketplace. No product launched.
Raptive / pay-per-value: Paul Bannister (Raptive CSO), via Digiday. Described pay-per-value as "the most complicated but by far the best model", explicitly compared to ASCAP/BMI royalties.
Really Simple Licensing 1.0: Specification finalised December 2025. Early backers: Reddit, Yahoo, Medium, O'Reilly, Ziff Davis. No major model developer publicly committed to honouring it as of early 2026.
x402: Released by Coinbase, May 2025. Contributed to foundation, April 2026. Cloudflare and AWS edge support; AWS Bedrock AgentCore payments; Visa Trusted Agent Protocol and Stripe agent-commerce integrations.
Cloudflare AI traffic data: Cloudflare "Agentic Internet" bot report, June 2026. 52% of AI crawler requests for training, up from ~22% spring 2025.
x402 adoption counter-evidence: Independent analysis (Chainalysis, x402stats.io, March-July 2026) found daily settled volume modest, with substantial test/non-organic activity and transaction-weight skewing toward larger payments over true micropayments.
Researcher citation preference: Documented across open-science and data-sharing survey literature; researchers consistently report reputational rewards (citations, recognition) as more important than direct financial compensation for contributing research outputs.


















%20Medium.png)


.png)