How this site was built
This site started in July 2026 as a place to keep the paperwork for one Mini Pro, and grew into a library of digital editions, a parts finder and a restoration log. It is one person’s project, built alongside the bench work. This page explains how it is put together — the editions, the hosting and the pipeline that publishes every change — for anyone who wonders how a site like this is made, or wants to build one for their own machine.
The short version↑ TOP
- A static website: every page is plain HTML built ahead of time. There is no server, database or login behind it.
- Hosted on Amazon Web Services — files in a private S3 bucket, delivered worldwide through the CloudFront content network.
- Every change is a commit to a Git repository. An automated pipeline tests it, builds the site and publishes it, usually within minutes.
- No advertising, no analytics scripts and no cookies of its own. Reader numbers on the Readers page come from the server's own logs.
- Built with heavy use of AI coding and research tools, directed and checked by me against the original paper.
Turning scans into digital editions↑ TOP
The heart of the site is the library: thirteen manuals, brochures and data sheets rebuilt as searchable, linkable digital editions, each with the original scan a click away.
A scanned manual is a stack of pictures of pages. Making one readable on a phone takes several passes:
- Split and clean the source. Each PDF is split into its sections without re-compressing the scans, so nothing is lost before the work starts.
- Transcribe the text. Pages are read with OCR — for the later editions, a vision model running locally on my own computer — and then checked page by page against the scan. Machine transcription is fluent and wrong in quiet ways: it silently corrects misprints, drops table rows and invents labels. Every page is compared with the original before it is published.
- Keep the paper's mistakes. These are historical documents, so the text says what Otari printed, misprints included. Where the printing is wrong, a visible editorial note marks the error and gives the correction, rather than quietly fixing it.
- Redraw the figures. Schematics, board layouts and exploded views are traced from the scan into sharp vector drawings, with the printed lettering recognised so the drawings can be searched. A tracing has to pass automatic checks — that its ink matches the scan, and its labels land where the printed ones are — before it is published. Next to every redrawn figure, a toggle shows the original scan, so you can always check the drawing against the paper. There are now around three hundred of them.
- Link everything. Figure and table references in the text jump to the figure; the parts lists feed the parts finder, which searches over three thousand factory part entries across the manuals; and cross-references point between editions that cover the same assembly.
The Mini Pro maintenance manual can also be downloaded as a Word document, generated from the same source as the web pages.
How the site is built↑ TOP
The site is written in Next.js and exported as static files. Pages are written in MDX — Markdown with the occasional interactive component, like the signal-path explorer on Inside the Audio. Tailwind CSS handles the styling.
Site search is Pagefind. It builds its index from the finished pages when the site is built, and the search itself runs entirely in your browser — no search query leaves your machine.
Around 270 content files and 150 automated test files sit in the repository. The tests check things a reader would notice when they break: links and anchors that resolve, figures that exist, parts entries that point to the right page, editions that match their sources, and redirects that still send old addresses to the right place.
Hosting↑ TOP
The finished site is a folder of files, and serving files is the cheapest, fastest and most secure thing a web host can do.
- Storage. The files sit in an Amazon S3 bucket that is private and encrypted. Nobody can read it directly — only the CloudFront distribution in front of it, using AWS's Origin Access Control.
- Delivery. CloudFront serves the site from edge locations around the world, over HTTPS only (TLS 1.2 or newer), compressed, and cached close to the reader.
- Addresses that keep working. A small CloudFront Function runs on every
request. It sends
wwwto the main address, maps folder addresses to their pages, and redirects the addresses of pages that have since moved, so old links and search results still land in the right place. - Domain and certificate. DNS is Amazon Route 53. The HTTPS certificate comes from AWS Certificate Manager and renews itself.
All of it is described as code — AWS CloudFormation templates kept in the same repository as the site — so the whole setup can be rebuilt from scratch, and every change to it is reviewed and recorded like any other.
The contact form↑ TOP
A static site cannot send email, so the contact form posts to a single small AWS Lambda function. It checks the message, verifies a Cloudflare Turnstile challenge (a privacy-friendly alternative to picture CAPTCHAs), and sends the message to me through Amazon SES. It runs only when someone writes in, so it costs nothing when idle. Its one secret is stored encrypted in AWS Systems Manager, never in the code.
Counting readers without tracking them↑ TOP
There is no Google Analytics or similar script on the site, and it sets no cookies of its own. The Readers page is built instead from CloudFront's own access logs. Once a day another Lambda function reads the logs, removes bots, crawlers and my own visits, groups the rest into visits, and publishes a small summary file that the Readers page draws its charts from. Countries come from the CloudFront location that served the page, not from looking up anyone's address. The method is described in full at the foot of that page.
Publishing a change↑ TOP
Every change, from a corrected part number to a new manual, goes through the same path:
- The change is committed to the Git repository, hosted on GitLab.
- GitLab CI installs the exact dependency versions recorded in the repository and runs the full test suite. If any test fails, nothing is published.
- The site is built, every internal link and anchor in the result is checked, and the search index is generated.
- The files are synchronised to S3. Scripts and styles have content-based file names, so they are uploaded first and cached for a year; the pages follow. Deploys run one at a time, so two quick changes cannot interleave and leave a page pointing at a file that has been removed.
- CloudFront's cache is cleared, and the change is live.
The pipeline signs in to AWS with short-lived credentials issued for each run, so no long-lived AWS keys are stored anywhere. Changes to the infrastructure, the contact form or the traffic counter deploy their own parts only when those files change; an ordinary content change only builds and publishes the site.
Built with AI, checked against the paper↑ TOP
Most of the code and much of the editorial groundwork was produced with AI tools — chiefly Anthropic’s Claude, working through planned, reviewed tasks, with local models doing bulk transcription. I set the direction, reviewed the output and checked it against the scans; the restoration and calibration on the bench were my own. The AI was fast and tireless, and it was also confidently wrong often enough that the checks described above exist because of it. The rule for the editions is simple: the scan is the authority, and anything that can be checked against it is.
Who built it↑ TOP
I work in cloud and AI engineering, and this site doubles as an example of that work. My consultancy, Expert Cloud & AI, has written up the project as a case study: From scanned manuals to a working restoration. If you have a similar problem — specialist documentation that people struggle to use, or a site that should cost less and break less — that is the place to get in touch.