How we work
Methodology
A result you cannot verify is not a result. This page describes where every claim comes from, what we do not claim and when we stop standing behind what we wrote.
Last updated: 21 August 2026
What evidence is
Most tools blend four different things into a single number. We keep them apart, because each one carries a different weight and different consequences.
| Term | What it means | Example |
|---|---|---|
| Observation | A raw fact with a source and a time. It does not judge. | The URL returned status 200 and a body containing a key. |
| Claim | Someone else's statement, usually from the vendor or a document. It has a quote and a location in the source. | “We do not use customer data to train models”, section 7.2 of the contract. |
| Finding | A confirmed problem backed by evidence that can be reproduced. | Account A read account B's data in the invoices table. |
| Risk | The impact of a finding in your specific use case. | Billing data of 34 agency clients exposed. |
| Decision | A verdict for a defined scope and date, with conditions. | Conditional, once the read policy is closed. |
Each of these carries a source, a time, an author, a confidence level, a scope and a date after which it is no longer current. Anything without a source and a date does not go into a decision, however plausible it sounds.
Authorization levels
Not every check can be run just because someone wants it. The deeper we go, the more consent it takes. The level is assigned to each test, not to the client, and without an active authorization the test does not run.
| Level | What is allowed | What it requires |
|---|---|---|
| L0 | What the system shows publicly: domain names, certificates, headers, published documents, public documentation. | Acceptance of the terms and a stated target. |
| L1 | Documents you hand over to us: the contract, the data processing agreement, an exported configuration. | Confirmation that you have the right to share them. |
| L2 | Read access to connected systems: the repository, the database schema, the cloud configuration. | Owner consent, an explicit permission scope, read-only access. |
| L3 | Tests that send real requests to a running system, including data boundary tests on test accounts. | A written authorization: asset, environment, time window, limits, emergency contact. |
| L4 | Repeated tests and monitoring of changes over time. | A standing authorization with a schedule and the ability to stop immediately. |
We do not bypass login, bot protection or rate limits on the system being checked. If something can only be established by getting around them, the result is “not established”, not “safe”.
How we measure confidence
We report four measures separately, because merging them into one number hides the most important piece of information: whether a result is low because the system is weak or because there was not enough data.
- Risk: how severe the impact is in your use case and how easy it is to trigger.
- Evidence strength: source quality, freshness, completeness, reproducibility and whether different sources agree.
- Exposure: who can reach this surface at all. Something reachable from the internet counts differently from something internal.
- Decision: the business outcome, meaning accept, accept with conditions, stop or not enough data.
Evidence strength decays over time. The same piece of evidence from a week ago and from a year ago does not weigh the same, and the report shows it.
What we do not check
This is the part most tools skip, and it decides whether a report can be shown to anyone.
No evidence does not mean “safe” and it does not mean “critical”. It means: we don't know. Every decision lists the areas we did not touch and says what it would take to close them.
Out of scope by default, unless we agreed otherwise:
- the production environment, unless there is a separate authorization;
- internal processes: who has access, how offboarding works, how passwords are stored;
- physical security and hardware;
- compliance with a specific standard or certification;
- legal assessment of contract terms, beyond pointing out what they mean for your use case;
- vulnerabilities disclosed after the date of our assessment.
Retests and expiry
We close a finding only when the fix passes the same test that detected it. A code change without confirmation closes nothing.
Every decision has a date after which it no longer applies. Not because the system suddenly gets worse on that day, but because after that point we have no basis to claim anything. A material change, such as a new subprocessor list or a change in access policy, reopens the decision earlier.
The vendor's right of reply
If an assessment concerns someone else's product, the vendor has the right to respond to the result before it goes anywhere beyond the person who requested it.
- The vendor sees the findings that concern them, together with the steps to reproduce.
- They can add their own statement. We publish it alongside the finding, not instead of it.
- They can report an error in a finding. We check it and correct it if they were right.
- They cannot remove a finding or change its content.
We do not publish lists of shame or security rankings of companies. A result applies to a scope, a date and a specific use case, so it does not belong in a table ranked from best to worst.
The role of the language model
We use language models to read documents, summarize, group and suggest questions. We do not use them to issue a verdict.
- The model suggests, a person approves. A suggestion is labeled as a suggestion.
- Every claim derived from a document has a quote and a location in the source.
- We treat the content of the document or page being checked as data, never as an instruction to the model.
- The model does not assign any status or change a decision.
- We do not send secrets or entire repositories to the model without your explicit consent.
Responsible disclosure
If we find a problem in someone else's product during an assessment, we tell the vendor first and give them time to fix it. Technical details go to no one except the client and the vendor.
If you found a problem in our system, describe it on the vulnerability disclosure page.