The Paper Giant
Archive
Paper Giant’s Archive makes past project work searchable and reusable. It brings reports, presentations, frameworks and other deliverables into one place where a team can find relevant examples, inspect individual pages and ask questions with references back to the source. This overview explains the experience and the system behind it for teams interested in a similar approach.
What teams can do
Find relevant past work.
Search by meaning as well as filename and project terms, then narrow results by project, client, sector, date, method or content type.Browse the work visually.
Explore page previews, search descriptions of diagrams and layouts, find similar pages, and open surrounding pages for context.Understand an engagement quickly.
Read document summaries covering purpose, methods, findings and recommendations, alongside project summaries and human corrections.Ask questions across the archive.
Request a short synthesis of relevant work with numbered source references. Connected AI assistants can also retrieve documents, page text and images through the same service.
For example, a team preparing a service design workshop could search for journey maps, inspect useful pages in context, read the associated project summaries and follow links to the original files. The aim is to make prior work usable without needing to remember its filename or who created it.
How documents become searchable
Mirror the source files.
Google Drive project folders are mirrored to network storage. The Archive reads that copy through a read-only mount, preserving the source files.Extract and organise.
Text is extracted from PDF, Word, PowerPoint and plain-text files, split into searchable passages and associated with project metadata. Filename and folder rules identify likely final versions and group related versions.Add descriptions.
Local models generate document summaries, project summaries, taxonomy labels and descriptions of rendered pages. Semantic and visual indexes support different ways of finding the material.Serve the results.
A browser interface, an HTTP API and an AI assistant connector expose documents and pages with source references. Scheduled jobs update the index; hashes and recorded progress allow processing to resume.
System architecture
| Component | Role |
|---|---|
| Network storage and application server | Holds the mirrored files, runs the FastAPI service and Qdrant search database, and stores processing state in SQLite. |
| Mac Studio | Runs the local language, embedding, reranking and vision models used to describe and retrieve material. |
| Browser and assistant interfaces | A React web app supports browsing and questions. A Model Context Protocol server lets compatible AI assistants use structured search, text and image tools. |
The core indexing and model processing run on the local network. Google Drive remains the upstream file source. Any connected external AI client has its own handling of the content it retrieves.
Source context and reuse
Results retain document and page references, with links to the archive and Google Drive where available. Page tools provide text, neighbouring pages, enlarged views and contact sheets so a reader can check an example before using it. Citations also travel with rendered images.
Search normally favours the latest non-draft deliverables and reduces duplicate results. Research transcripts, raw notes and contract paperwork are excluded from ordinary document and page discovery unless explicitly requested. Human-maintained project records can correct generated descriptions and set sharing labels without those edits being overwritten by regeneration.
Access and practical limits
The deployment is intended for a trusted office network and private remote access. Read access is available to anyone who can reach the service; processing operations require a key, and saved collections and research tracking use named client tokens. This is a network access model, rather than enforcement of each person’s Google Drive permissions.
Material defaults to internal precedent use. Its sharing guidance calls for removing identifying client and participant details before external use; documents marked restricted do not serve page renders or text. Finding a document does not itself establish permission to distribute it.
AI summaries, page descriptions and version classifications can be wrong or incomplete. Answers are instructed to use retrieved excerpts and cite their claims, but the original documents remain the basis for checking an answer. Coverage depends on what has been mirrored and successfully processed; a search result is not a complete inventory.
A useful first demonstration
Use an approved sample of documents to follow one real question from search to a page preview, then to a cited answer and the original source. That shows both the discovery experience and the evidence trail. A team exploring its own deployment would need a source collection, project metadata, model capacity and access rules suited to its audience.