Skip to main content
Smartbox.ai

Technology

Built so it cannot tell you what you want to hear.

“Do you use AI?” is rarely the real question. The real ones are where your models were trained, whether our data trains them, whether you can explain a decision, and whether you can prove what came out. Those are the questions this page answers.

How it works

Nine things that make a disclosure defensible.

  • Your data is not our training data

    We do not train on customer material. Our models are trained on our own content, on synthetic data we generate, and on sources we hold written authorisation to use — with a named owner and a date against that authorisation. Nothing you put into Smartbox teaches the model, and nothing another customer puts in has taught the one that reads your files.

    Ask any AI vendor where their training data came from. It is the question with the fewest good answers, and it is the one your DPIA will eventually have to record.

  • Every part of the model has to justify itself, in writing

    Each component we build on carries a written statement of what it is, where it came from, and why it is defensible — reviewed and signed. Anything that cannot clear that bar does not ship. The one publicly available component we use is openly licensed, its own training data is declared and public, and it is used only while we train: it never reaches the system that reads your documents.

    This is the paperwork nobody enjoys producing and everybody wants to see when a decision is challenged. We would rather have it already written than assemble it under pressure.

  • We can show you why, not just what

    When the platform flags something as personal data, it can show which words in the document drove that decision — not just the verdict. Confidence scores are calibrated, so a stated confidence reflects how often that kind of judgement actually turns out to be right, rather than being a number the model prints about itself.

    “The system flagged it” is not an answer anyone can defend. “It flagged it because of these words, and here is how reliable that has been” is.

  • The same file, the same answer, every time

    Run the same document through Smartbox twice and you get the same result. Two reviewers, a month apart, get the same result. Every version of our detection is fingerprinted, so if a decision is questioned a year later we can identify exactly which one made it and run the file through it again.

    This is what a Tribunal question actually sounds like: not “is your AI accurate?” but “can you show me it would do the same thing again?”

  • We check the redaction actually worked

    Most tools redact and move on. Smartbox goes back and reads its own output: the finalised copy is re-scanned and re-analysed, and the result is compared against every value you decided to redact. If something survived, you are told before the bundle leaves the building — not after the requester finds it.

    This is harder than it sounds, which is why it is rare. The disclosure copy is a flattened image with no text underneath, so it has to be read back optically. And a name split across two lines will not be found by searching for the name — so the check has to be more sensitive than the redaction that preceded it, then risk-ranked for a human rather than waved through.

  • Three answers, not two

    Most software has two answers: fine, or not fine. Smartbox has a third — we could not establish either— and it never reports that as the first. A check that timed out reads “couldn’t verify”, not “clean”. A check that ran with nothing to look for is reported as having verified nothing, because zero things to search for means nothing was verified, not that nothing was wrong.

    That second one is the failure this was built to kill: a box-level check that cheerfully reports success when it had no values to hunt for. And a control switched off in your environment is written onto the response, so an operator reading a success can see which checks did not happen — the difference between “this passed” and “nothing asked”.

  • Redactions are burned in, not drawn on

    The two ways organisations leak redacted material are a black rectangle drawn over live text in a PDF, and a marker pen that photocopies through. Smartbox produces flattened images: the underlying content is gone from the file, not hidden behind something. There is no layer to move aside and nothing to recover with a text selection.
  • Audio and video, not just documents

    Recorded calls, interview tapes, bodycam footage and dictation go through the same pipeline as documents. Audio is transcribed with speaker separation and automatic language detection; video has its audio track demuxed and treated the same way. Because a transcript is text, the same detectors that find personal data in a document find it in a recording.
  • The interface never shows you a number it cannot source

    A rule we build to: no mock data, no sample rows, no invented counts, anywhere in the product. Every screen either shows real data from a real source or says plainly that it has none. It is enforced in code — an unimplemented data source raises an error rather than quietly returning something plausible.

    Invented data makes a screen look like it works. For a team whose job is to prove what happened, a screen that looks like it works is worse than one that admits it does not.

See it on your own files.

Thirty minutes with a product expert — your data or ours.

Book a demo

We use cookies to measure how the site is used, and — if you agree — to measure our advertising. You can accept one without the other, and declining is one click. See our cookie policy.