Model integrity
LLM Fingerprinting: How to Verify Which Model an API Serves
You pay for a flagship model. A provider can serve you a cheaper one and bill you the same. Here is how to catch it, from a three-call tokenizer check to hardware proof.
The model name is a string, not a receipt
Every API call carries a model name. You type it, the provider echoes it back, and your invoice charges for it. None of that tells you which weights did the work.
The incentive to cheat is large. A frontier model can cost many times more to serve than a small open model or a 4-bit copy of itself. Route even part of the traffic to the cheap option, keep charging full price, and the difference is margin. From your side, the response looks like any other response.
This already happens. In 2024 a Stanford team ran Model Equality Testing against 31 commercial endpoints selling four Llama models. Eleven served output that did not match the weights Meta published. In 2026 a CISPA-led audit of shadow APIs, the unofficial resellers that sell cheap access to frontier models, found identity checks failing in 45.83% of fingerprint tests and task scores up to 47.21% below the official API. These were not fringe services. Researchers had used 17 of them in 187 published papers.
The resale market is worse. A knowledge-boundary audit of six shadow API platforms found 7 of 27 platform and model pairs statistically inconsistent with the official endpoints, mostly on the premium models people pay the most for. Another study found a flagship-branded endpoint on a large aggregator that was indistinguishable from an open-weight model.
If you build on a hosted model, you need a way to check. What follows is the toolkit, ordered from cheapest to strongest, with what each method catches, where it breaks, and code you can run today.
Five ways you get a different model
Substitution is rarely a dramatic swap of one brand for another. The common versions are quieter, and the quiet ones are the hard ones to catch.
Every call goes to a different, cheaper model.
Easy to catchThe same model at lower precision. Same habits, slightly worse answers.
Hard to catchAn older or smaller member of the family you asked for.
MediumOnly some calls are downgraded, so a single spot check passes.
Hard to catchSame weights, but a hidden prompt, a trimmed context or a filter in front.
MediumQuantization and dilution pass most one-off tests. Both need statistics over many calls.
A full swap is the easiest to spot, because different model families tokenize differently, know different things and write differently. Quantization is the hardest. The same model at lower precision keeps its habits and gets a little worse, and that drift can hide inside normal sampling noise.
Dilution is the clever one. Route 30% of calls to the cheap model and any single spot check will probably see the model you paid for. IRIS, an auditor built for exactly this case, caught 30% dilution with 0.85 power at a 1.7% false-positive rate, and estimated the swapped fraction to within 4 points.
Wrapper changes leave the weights alone. A relay can inject its own system prompt, cut your context window to save memory or put a cheap filter in front. The open-source api-relay-audit tool tests for these: it compares token counts to find hidden prompts and plants canary markers through a long prompt to find where the context gets cut.
None of this needs bad intent. Providers reroute under load and swap inference kernels. IRIS flagged 14 of 15 pairs of providers serving the same model, purely from quantization and kernel differences. A study of ten commercial gateways found silent model switches, memory that degraded across turns and prices that drifted from the published rate. Whatever the reason, you are not getting what the label says.
Do not ask the model who it is
The first test everyone tries is "Which model are you?" It is worthless.
A model's answer about itself is generated text like any other. It comes from the training data and the system prompt, and the provider writes the system prompt. One line, "You are Flagship Pro", and a small open model will say it is Flagship Pro all day. It fails the other way too. With no prompt at all, models often name the wrong vendor, because their training data is full of other models' output.
Its tokenizer, number habits and speed all match the small model. The name came from one line it was told to say.
A forensic protocol for anonymous models puts it bluntly: self-identification is untrustworthy by design. Throw the answer away and measure behaviour instead.
The audit ladder
No single cheap test proves which model you are talking to. There is a ladder. Each rung costs more and tells you more. Start at the bottom, because the bottom rungs end most investigations.
The ladder also tells you what to expect. A full swap usually fails on the first two rungs. A 4-bit copy of the right model can pass everything up to the distribution tests. Partial routing can pass any test you run once.
Rung one: tokenizer, context and cutoff
Start with signals that cost almost nothing and are hard to fake without changing the model.
Token counts. Most OpenAI-compatible APIs report prompt_tokens in the usage block. Send the same text to the endpoint and to a provider you trust. Models that share a tokenizer and chat template return the same count, give or take a fixed offset for the template. A different count means a different tokenizer, and that rules the claimed model out.
Fingerprinting a relay costs pennies.usage.prompt_tokensexpected 8returned 12count() { # count <base-url> <api-key>
curl -s "$1/v1/chat/completions" \
-H "Authorization: Bearer $2" -H "Content-Type: application/json" \
-d '{"model":"'"$MODEL"'","max_tokens":1,
"messages":[{"role":"user","content":"Fingerprinting a relay costs pennies."}]}' \
| jq .usage.prompt_tokens
}
count "$OFFICIAL_URL" "$OFFICIAL_KEY" # the provider you trust
count "$ENDPOINT_URL" "$ENDPOINT_KEY" # the endpoint you are auditingBe careful with a match. A holdout study of token-count fingerprinting found the signal reliably identifies a shared tokenization stack but not a shared model. Some same-family pairs fell below its threshold, and 157 successful responses came back with no usage data at all. Use counts to rule models out, never to confirm one.
Context length. Plant unique markers through a long prompt and ask for them back. Binary-search the length where recall breaks. If a model sold with a 200k window starts losing markers at 32k, something in the middle is trimming your context.
Knowledge cutoff. Every model has an edge to what it knows, and Dated Data showed the effective edge often sits earlier than the published date. Ask about dated events month by month and you get a rough age for the weights. KBF made this rigorous by probing numerical facts near that edge. Across 16 production endpoints it flagged all 155 economically relevant substitutions without rejecting a single same-model control, and caught mixed routing with only 5 to 10% of traffic swapped.
Rung two: number habits and probe batteries
Ask a model to pick a random number between 1 and 100 and it will not. Models have favourites, and the favourites differ by model.
One Token Is Enough turns this into a fingerprint. It asks trivial one-word questions like that in four languages, takes one output token per answer, and compares the spread of answers with Jensen-Shannon divergence. Across 165 models on a large aggregator, two halves of one model's samples sat an order of magnitude closer than samples from different models. With about a hundred single-token queries the verification error rate was below 11%, and 7.3% with the full battery. IRIS uses the same idea, asking for random numbers and strings, and told apart closely related models in one family at 0.99 AUROC.
Pick a random number between 1 and 100.Here is the whole test. Sample the trusted provider twice to measure the noise floor, then compare it with the endpoint. A gap many times the noise floor is a strong signal.
from collections import Counter
from openai import OpenAI
from scipy.spatial.distance import jensenshannon
PROMPT = "Pick a random number between 1 and 100. Reply with the number only."
def profile(client, model, n=100):
picks = Counter()
for _ in range(n):
r = client.chat.completions.create(
model=model, temperature=1, max_tokens=4,
messages=[{"role": "user", "content": PROMPT}])
picks[r.choices[0].message.content.strip()] += 1
return picks
def divergence(a, b):
keys = sorted(set(a) | set(b))
return jensenshannon([a[k] for k in keys], [b[k] for k in keys]) ** 2
official = OpenAI(base_url=OFFICIAL_URL, api_key=OFFICIAL_KEY)
endpoint = OpenAI(base_url=ENDPOINT_URL, api_key=ENDPOINT_KEY)
baseline = divergence(profile(official, MODEL), profile(official, MODEL))
suspect = divergence(profile(official, MODEL), profile(endpoint, MODEL))
print(f"noise floor {baseline:.3f}, endpoint {suspect:.3f}")Richer prompts give richer signal. LLMmap picks a handful of thematically varied questions where models answer in telling ways, then classifies the answers. With as few as eight queries it identified 42 model versions with over 95% accuracy, even behind unknown system prompts, retrieval and chain-of-thought wrappers.
It works because models have idiosyncrasies. A classifier trained on raw outputs told five major chat models apart with 97.1% accuracy, and the signal survived when another model rewrote, translated or summarized the text. Word choice is a fingerprint.
Rung three: is it the exact weights?
Probes tell you the family. They usually cannot tell you whether you are getting full precision or a 4-bit copy, because a quantized model keeps its personality. For that you need statistics.
Treat it as a two-sample test. Sample many completions from the endpoint and from a reference you trust, ideally the open weights on your own hardware, then ask whether both came from the same distribution.
Model Equality Testing used Maximum Mean Discrepancy with a simple string kernel and reached a median 77.4% power across a range of distortions, with about ten samples per prompt. The Rank-Based Uniformity Test does better on a tight budget. It ranks the endpoint's answers against the local model's own samples, avoids query patterns a provider could spot, and beat the kernel test at catching a model mixed with its own 4-bit version.
The catch: closed models have no public weights, so you compare against the official API instead. That only works when the official API is not the thing you suspect.
Rung four: log-probabilities
If the endpoint returns log-probabilities, you can look under the hood.
Most models produce their output through a final layer that is much narrower than the vocabulary. Finlayson and colleagues showed this softmax bottleneck pins every output to a small subspace, so a single full output can identify the source model and an audit can spot quiet updates. They estimated the hidden size of a major commercial model at about 4,096 for under $1,000 of queries. A DeepMind-led team recovered the full projection layer of two production models for under $20.
There is a simpler version for open models. Strake first tried grading answers by quality, and it failed backwards: the cheaper impostor wrote simpler, more predictable text, which scored better on 9 of 10 prompts. What worked was a scoring challenge. Fix a token sequence and ask the endpoint how likely it is. The advertised model is more confident in its own text. A 6x cheaper impostor lost 10 of 10 challenges. A 2x cheaper one lost 8 of 10, which sits close to ordinary quantization noise, so it took 10 to 15 challenges to separate the two.
That is the rule for subtle swaps: one check is noise, a running average is a verdict. It is also why Cai and colleagues call software-only detection unreliable. Production inference is not deterministic, so single log-probability readings wobble, and small substitutions hide inside the wobble.
Rung five: the clock
Even when the API returns nothing but text, it cannot hide how fast the text arrives. Time to first token and tokens per second depend on model size, precision, hardware and serving tricks.
LeakyLMs modelled token timing on modern GPUs and used it to detect speculative decoding in a production API and to recover architecture details. For Llama models, the near-correct configuration was in its top ten guesses more than 90% of the time. Carlini and Nasr showed timing leaks even through encrypted traffic.
An audit needs much less precision than that. Measure the advertised model where you trust the source, and the likely substitutes. If your endpoint keeps running at small-model speed, that is a strong tell. Network jitter is the enemy, so take many samples and compare medians, not single calls.
When the provider fights back
Everything above assumes the provider is not watching for your audit. Assume it is.
A provider that wants to pass audits can send anything that looks like a known probe to the advertised model and everything else to the cheap one. It can downgrade only some of the time. It can even train the cheap model to pass. GhostPrint did exactly that: it fine-tuned weak models to mimic a stronger one and consistently got past representative fingerprinting methods while keeping the weak model's low cost. The authors proved that a finite query budget and a weak classifier leave every user-side audit open to this.
That does not make auditing pointless. It changes how you do it.
- Hide probes in normal traffic. Not in a burst at 9am every Monday.
- Rotate them. A fixed probe set is a pattern a router can learn.
- Audit all the time. Dilution only shows up across many requests.
- Stack signals. Faking token counts, knowledge, number habits and timing at once is much harder than faking one.
The only proof: hardware attestation
Every method above infers the model from outside. Attestation replaces inference with proof.
Isolated from the host
enclave: sealedHash of weights and code
sha256: 9f2c…a71eSigned by the hardware key
sig: validCheck against vendor keys
model: flagshipIn a trusted execution environment, the model runs in an isolated enclave and the hardware signs a measurement of exactly what it loaded. The response carries that signed receipt, and you check it against the hardware vendor's keys before you trust the output. Cai and colleagues tested this and found it gives a cryptographic guarantee of model integrity for a modest performance cost. The format is being standardized too, in an IETF draft for an Attested Inference Receipt.
The limit is plain. The provider has to offer it. Until yours does, you are building a case, not proving one.
An audit you can run every week
- Write down what you expect. Model, version, context length, tokenizer, published cutoff. If the weights are open, run them yourself as the reference.
- Run the cheap checks once. Token counts, context canaries and a cutoff probe. Many bad endpoints fail here in under fifty calls.
- Fingerprint the family. A hundred random-number queries and a short probe battery, compared with your reference and its noise floor.
- Track drift. Run a distribution test on a fixed prompt set each week and log timing medians.
- Hide and rotate. Mix probes into production traffic at random times and change them each cycle.
- Confirm before you accuse. Rerun a hit with a bigger sample, then ask the provider for attestation.
Here is the whole toolkit on one page.
| Method | Needs | Catches | Limit |
|---|---|---|---|
| Tokenizer and context | Text API | Wrong family, trimmed context | Proves a shared tokenizer, not a shared model |
| Knowledge boundary | Text API | Older versions, mixed routing | Needs a fresh probe set as models update |
| Number and probe fingerprints | About 100 calls | Which family and version | A tuned impostor can copy the habits |
| Distribution test | A trusted reference | Quantization and fine-tuning | Noisy on short answers |
| Log-probabilities | Logprob access | Source model, quiet updates | Inference noise blurs small swaps |
| Timing | Many samples | Size, precision, serving tricks | Network jitter |
| Attestation | Provider support | The exact weights and code | Only if the provider offers it |
If you sell the model: the same tools catch resellers
Turn the problem around. If you are a lab or an inference host, people resell your model through relays built on stolen keys and farmed accounts. The fingerprinting a buyer uses to audit you is also how you prove a relay is selling your model.
Buy access like any customer, send fingerprint probes and compare the answers with your own model. A match is evidence. Pair it with the accounts the traffic ran through and you know whose access to cut.
Agents search chat channels, marketplaces and storefronts.
Like any other customer.
Match them against your model.
Match: your modelTrace the accounts and stop the next sign-up.
Blocked at registrationThat is how Relay Watch works. Our agents search the resale market, buy access to relays, fingerprint the answers against your platform and trace each match to the accounts behind it. NoResponse then blocks new sign-ups from the same operation at registration. If you sell model access, request a threat briefing and we will show you which relays sell your tokens today.
Frequently asked questions
How do I check which model an API is really serving?
Do not ask it. Check what is hard to fake: the prompt-token count, how much context it keeps, what it knows, and how it answers a fixed set of probes such as picking a random number. For quantized or partly rerouted models, add a distribution test against a trusted reference and repeat it over time.
Why can't I just ask the model what it is?
Because the answer is generated text. A hidden system prompt can make any model claim any name, and models often name the wrong vendor even with no prompt. Researchers who audit anonymous models treat self-identification as untrustworthy by design.
What is model substitution?
It is when an API serves a different model from the one it charges for: a cheaper model, a quantized copy, an older version, or a mix where only some calls are downgraded. A 2024 audit found 11 of 31 commercial endpoints for four Llama models serving output that did not match the published weights.
What is an LLM relay or shadow API?
A third-party service that resells access to models, usually at a steep discount. Some pass calls to the official API, some run on stolen keys or farmed accounts, and some quietly serve a cheaper model. An audit of shadow APIs used in research found identity checks failing in 45.83% of fingerprint tests.
Can you detect a quantized model behind an API?
Yes, but not in one call. A quantized copy keeps the model's habits, so you need a two-sample test that compares many answers from the API with answers from the reference weights. Kernel tests work with about ten samples per prompt, and rank-based tests do better on small budgets.
Does the token count tell me which model I am using?
It tells you which tokenizer you are using. A different count rules a model family out. A matching count only shows a shared tokenizer, so use it to exclude models, never to confirm one.
Can a provider fake a model fingerprint?
Yes. Researchers fine-tuned weak models to pass fingerprint checks built for a stronger one. Hide your probes in normal traffic, rotate them, combine several signals and audit continuously. Only hardware attestation proves which model ran.
Sources
- Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs Cai, Shi, Zhao, Song, 2025.
- Model Equality Testing: Which Model Is This API Serving? Gao, Liang, Guestrin, ICLR 2025.
- Real Money, Fake Models: Deceptive Model Claims in Shadow APIs Zhang et al., 2026.
- KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing Fang et al., 2026.
- One Token Is Enough: Fingerprinting and Verifying LLMs from Single-Token Output Distributions Bruckner, 2026.
- IRIS: Budgeted Black-Box Auditing of Model Substitution and Routing Dilution in LLM Gateways Zhang, Zhang, Qin, 2026.
- Behavioral Consistency and Transparency Analysis on LLM API Gateways Lin et al., 2026.
- api-relay-audit: security audit tool for third-party AI API relays Open-source tool, 2026.
- Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification Xi, 2026.
- Token Counts Are Not Model Lineage: A Frozen-Threshold Holdout Study Chen, 2026.
- Dated Data: Tracing Knowledge Cutoffs in Large Language Models Cheng et al., 2024.
- LLMmap: Fingerprinting for Large Language Models Pasquini, Kornaropoulos, Ateniese, USENIX Security 2025.
- Idiosyncrasies in Large Language Models Sun et al., ICML 2025.
- Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test Zhu et al., 2025.
- Logits of API-Protected LLMs Leak Proprietary Information Finlayson, Ren, Swayamdipta, 2024.
- Stealing Part of a Production Language Model Carlini et al., 2024.
- Can You Tell When an LLM API Swaps in a Cheaper Model? Strake, 2026.
- Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing Majidi, Mireshghallah, Taram, 2026.
- Remote Timing Attacks on Efficient Language Model Inference Carlini, Nasr, 2024.
- Your "Pro" LLM Subscription May Actually Be "Free": Fingerprint Spoofing Risks in LLM Inference Services Zhang, Li, Wang, 2026.
- Attested Inference Receipt (AIR): A COSE/CWT Profile for Confidential AI Inference IETF Internet-Draft, 2026.
Find out who resells your tokens
If you sell model access, we search the resale relays we monitor for your platform and show you which sign-ups we would have blocked.
Request a threat briefing