Media
Your Archive Is the Asset, and Most of It Is Unreachable
By Vivek S N · 2 July 2026 · 3 min read

Photograph by Sear Greyson on Unsplash
Ask a publisher what their most valuable asset is and they will say the archive. Look at where the engineering budget goes and it is almost entirely on producing and publishing new material.
This is not irrational — new content drives subscriptions and attention. But it means that in most publishing organisations the largest body of owned, paid-for, rights-cleared material is reachable only by someone who already knows it exists and roughly what it was called.
That gap has become considerably more expensive in the last two years.
Keyword search was never adequate, and now it shows
Search over a large archive has always underperformed. The vocabulary shifts across decades, terminology in any specialist field changes, and the words a reader uses today frequently do not appear in an article from 2009 that answers their question precisely.
Editors have compensated by carrying the archive in their heads. That works until they retire.
Semantic retrieval closes most of this gap. An article about a concept under its old name becomes findable by its current name, and a question phrased as a question finds material written as an explanation. For a large archive this is often the single highest-return piece of engineering available, because the content already exists and is already paid for.
The archive is also now a licensing question
Model developers need high-quality, rights-cleared text. Publishers own exactly that.
The organisations doing well from this are the ones whose archives are structured, deduplicated, clearly rights-attributed and queryable through an interface — able to license access as a product. The ones doing badly have a valuable archive locked in a content management system that cannot express what it holds, and are discovering it was scraped anyway.
Getting from one position to the other is mostly unglamorous data work: normalising formats accumulated over decades, resolving where rights actually sit, deduplicating the same article syndicated four ways, and building retrieval that can answer questions about the collection rather than return a list from it.
Generated summaries and editorial authority
There is obvious appetite for AI summaries over an archive, and obvious risk. A publisher's credibility is the product, and it is not worth an efficiency gain.
Two rules make this workable. Generated text is labelled as generated — not in a footer, at the point of reading. And every generated statement remains traceable to the source article, so a reader can go and check, and so an editor can audit what the system is claiming on the masthead's behalf.
A summary that cannot be traced to a source should not be published. That constraint sounds restrictive and mostly just rules out the applications that were going to cause trouble.
Entitlement is a retrieval problem now
Once retrieval spans the archive, access control has to move with it. An institutional subscriber entitled to one collection but not another must not receive an answer synthesised from both — and the check has to happen inside retrieval rather than at the article boundary, because a generated answer has no article boundary.
Getting this wrong in the permissive direction is a licensing breach. Getting it wrong in the restrictive direction blocks a paying institution, which costs more than the leakage it prevented. Both need to be tested against real entitlement data before launch, not modelled optimistically.
Do not break the citations
Any archive project eventually proposes a replatform, and the risk that gets underestimated every time is link continuity.
For a publisher, inbound links and citations are accumulated authority built over years, and in academic and clinical publishing they are also part of the scholarly record. A migration that breaks them discards something that cannot be repurchased.
URL mapping and redirects belong in the requirements, planned before migration and verified afterwards against real traffic and real citation data. This is tedious and it is not optional.
iLeaf has built content and evidence platforms including BMJ's digital transformation, interactive digital publications, and retrieval over large archives — see Media & Publishing.
Thinking about this for your own business?
We have been building and running enterprise systems since 2011. Talk to a solutions lead about where agents pay off first.
Talk to a solutions lead