What this cannot do
If you read one page before buying, read this one. Nothing here is legally required boilerplate; it is the list of things we know are wrong or unproven with our own product.
Last checked against our measurement record: 14 August 2026. Every number on this page is dated where it appears, because a measurement with no date attached is how a stale figure survives a re-measurement — which is exactly what happened to the extractor number below, and it sat here wrong for two weeks.
It is not a diagnosis, and it is not close to one
It cannot tell you whether you have any condition. It will not name one. It has no clinical validation, no clinician in the loop, and no cohort behind it.
The report describes patterns in text you wrote. Turning a pattern into a meaning about a person is a job for a person — ideally one who is qualified and who knows you.
If something in the report worries you, the useful next step is a conversation with a professional. It is not a second report.
We have run this on one person
N = 1. One real ChatGPT export has been through the full pipeline end to end.
The central question this project was built to answer is: does it tell two people apart, or does it flag everyone the same way? We cannot answer that yet. Answering it requires at least two subjects, and we have one.
That is not a small caveat. A tool that produces a plausible, well-cited, internally consistent report about everyone who uses it would look exactly like this one does today. We do not think that is what we have — but we have not proved it isn't, and we would rather sell you a report with that written on the box.
The numbers we measured, including the bad one
Run the same corpus through twice and you do not get identical output — the models are not deterministic, and chasing determinism would cost more than it buys. So we measure agreement between runs instead. All three figures below were measured in July 2026 on the one corpus we have, and none has been re-measured since.
| What | Agreement | Read this as |
|---|---|---|
| Which sessions are notable | 0.78 | Fairly stable. Two runs mostly point at the same places in your history. |
| Which pattern type a session gets | 0.50 | Weak. Two runs agree about where to look far more than about what to call it. |
| The synthesis stage | 0.83 | Stable, after we found and fixed a structural bug that had it at 0.11. |
The 0.50 is the honest problem on this page. It means the label attached to a stretch of your history is roughly a coin-flip between runs, even though the location is reliable. This is why the report gives you the quotes, the everyday explanation, and the counter-argument rather than a confident category — the underlying material is more trustworthy than any name we put on it.
Our extractor has a known blind spot
When we tested the extraction model against hand-built cases, it found 81% of what it was looking for — 39 of 48 trials — and stayed correctly silent on all 56 control cases that contained nothing to find. That is the figure for the model and the prompt that will actually run on your export, measured on 2026-07-25.
Underneath that average, recall is not even across pattern types. On an earlier version of the prompt we measured one type at 0 out of 3: how a person responds when the assistant contradicts them. Those were three trials on a prompt we have since changed, so the number is stale in both directions — we do not claim it is still 0 out of 3, and we cannot claim it is fixed. We have not re-measured recall per type on the current prompt.
That matters more than the aggregate suggests, because the report treats the mix of pattern types as meaningful. An extractor that cannot see a type produces a profile where that type is nearly absent, and a reader interprets absence as a fact: "this person doesn't dig in when challenged." The tool's limitation gets rendered as a statement about a human being.
Here is what we actually do about that, and it is less than we would like. Your report names the extraction model, prints the aggregate recall above, and states that anything the extractor cannot see is missing from the report — that absence there means "not detected" and never "not present". That is a general warning.
What it is not is the specific safeguard this section used to promise. Until 2026-08-14 this page said that any pattern type below the recall floor is named in your report and that the corpus-level pass is forbidden from making claims about it. Neither of those is built. The per-type breakdown is not printed, and nothing stops the corpus-level pass from drawing a conclusion about a type the extractor may be weak at. It was written here as a description and it was really a plan; our own internal record had it correctly filed as unfinished the whole time, which is worse rather than better. It is fixed here by deleting the claim, not by shipping the feature — the feature needs a per-type re-measurement we have not paid for yet.
The principle we hold to in choosing models is that uniform mediocre recall is safer than lopsided excellent recall. In practice that has not yet been the deciding factor in a binding: the current extractor was chosen on the aggregate gate and on control failures — the candidate we rejected was rejected for making false claims on 2 of 24 control cases, which is the failure that matters most, not for having a lumpy profile.
A guardrail that leaks about one time in six
The report is instructed never to use clinical or diagnostic register about you. We measured how well that instruction holds by generating the same report repeatedly and having independent judges check the text.
It leaks in roughly 1 of 6 draws. Not zero. An instruction-based guardrail is probabilistic and fails silently — which is why the instruction is no longer the only thing standing between you and a report that talks about you in diagnostic language.
Two things now sit behind it. A fixed word list is checked in code against every sentence the model writes about you, and the report is refused outright if it hits — no model judgement involved, so it cannot have an off day. Words you used about yourself are deliberately exempt: if you wrote "I feel manic", that is your word about your own life, and hiding it would conceal what we were looking at rather than protect you. And a panel of four independent judges from different vendors reads the finished report as a structural gate. Neither is a guarantee. The word list only knows the words on it.
What we would need to see to stop selling this
These were written down before we had results, so that a disappointing outcome could not be quietly re-interpreted as a good one. We publish them because a claim about your own falsifiability is worthless if you keep it private.
We stop, and write up the negative result, if:
- It does not discriminate — subjects get effectively identical profiles, and tightening the pipeline does not change that.
- Findings cannot be traced — a large share of candidate findings fail the citation gate, meaning the model is inventing them.
- Reproducibility collapses — agreement stays below 0.5 after addressing the cause.
- It causes harm without being useful — a subject reports the output was distressing and told them nothing, or a report reads as a de-facto diagnosis despite the guardrails.
- Consent or privacy fails — anyone withdraws, or the redaction pass is found leaking personal information.
Number 3 came close. Synthesis agreement measured 0.11 at first. The cause turned out to be structural rather than a prompt problem, and the fix took it to 0.83 — but it was a genuine near-miss, and the kill-switch was worded in a way that would have pointed at the wrong remedy.
Things it structurally cannot see
- Anything you did not type. Voice conversations, other assistants, other apps, and your entire life away from a keyboard are all invisible. A chat log is a keyhole.
- Why. It sees that a topic returned eleven times. It has no idea whether that is rumination, a research project, or your job.
- Context that changes everything. A bereavement, a deadline, a new diagnosis, a move — the export contains only what you happened to type near those events.
- Selection bias in the medium itself. People bring problems to an assistant and solutions to nobody. Chat logs are systematically skewed toward unfinished and unresolved things. The report is reading a biased sample and cannot correct for it.
What the research actually shows
In May 2026, ETH Zurich researchers published the first substantial study on this: 62,090 real ChatGPT conversations donated by 668 people, who also sat a standard personality questionnaire so there was something to check against.
The result is genuinely encouraging for the premise and genuinely modest in size. A model trained on that data could sort people into low, medium or high on each of the Big Five traits better than chance on all five — reaching about 61% accuracy where guessing gets 33%. Some traits were far easier to detect than others.
Two things follow, and we would rather say both.
There is real signal in chat history. That was not established before, and it is the reason a tool like this is worth building at all.
It is much less than the marketing in this category implies — including claims you may have seen that this kind of analysis is more accurate than a personality questionnaire. In that study the questionnaire was the answer key. And what was demonstrated was a coarse three-way sort by a model specifically trained for it, not a language model reading your history and telling you your attachment style, your blind spots, or a score out of a hundred.
It is also not evidence for this tool. That study tested trait inference, which is the thing we deliberately do not do. We would be borrowing credibility we have not earned. What we take from it instead is the finding that detectability varies with what people happened to talk about — which is the bias described immediately above, confirmed by someone other than us.
You were not the only author of your own messages
This is the limitation we are least able to do anything about, and it has become more serious rather than less.
Every message you wrote was a reply. It was shaped by what the assistant had just said — its register, its enthusiasm, how much it agreed with you, whether it asked a follow-up. People mirror their interlocutor, and an assistant tuned to be warm and agreeable is a strong one to mirror. So the transcript is not a record of you thinking. It is a record of you thinking with that particular thing, in that particular period.
Assistant memory tightens this into a loop. Once a system synthesises a profile from your earlier messages and conditions later replies on it, the sequence becomes:
- You write the kinds of things people write to an assistant — stuck, curious, at 1am.
- It remembers that slice and treats it as who you are.
- Its replies are shaped by that slice.
- You respond to those replies, writing more of the same kind of thing.
Memory did not fix the sampling problem. It industrialised it. What we read as your pattern may be a groove the two of you wore together.
Three consequences we want you to hold while reading your report:
- A pattern we report may be an artifact of the loop rather than a fact about you. We cannot separate the two from your side of the transcript alone, and we do not pretend to.
- Your corpus is not a stable object. Two exports from the same person at different times sit on different memory states and different model versions. Some of what looks like change over time is the assistant changing, not you.
- This runs deeper than our own exclusion rule. We bar the model that wrote your assistant turns from judging them. But through memory and mirroring, that model also shaped your user turns — so the corpus is more entangled with it than "don't let it grade its own homework" quite covers.
There is a second pipeline in this project that reads the assistant's turns instead of yours, precisely because the human-side screener cannot see this. It produces population-level metrics and no document about any individual. It is built and almost entirely unmeasured, so nothing in your report comes from it — but that is where an honest answer to this would have to come from.
If reading it is hard
The report quotes you back to yourself, and that can land harder than expected — particularly if your history contains a bad period. If your export includes messages about self-harm or ending things, the report will say so before you reach them, because being surprised by your own worst night is worse than being warned.
If any of it is heavy, go through it with someone you trust. If you are in crisis right now, please use one of these — this document is not a substitute for a person.