Black-box model audit

One token leaves a behavioral trace.

Compare a trusted deployment with an API under test using the 2026 PAMELA probe battery: one-word answers, multilingual distributions, and Jensen–Shannon distance.

240 total API calls
Trusted enrollment

Reference API

REF
Endpoint under test

Target API

TGT

Keys are stored only in this tab's session so interrupted audits can resume. They are cleared when the tab closes and are never included in exports.

How it works

Compare response distributions, not a single answer.

Each fixed one-answer probe is sent 15 times to both the trusted reference and the endpoint under test. The response counts become two empirical probability distributions. The tool calculates their Jensen–Shannon distance for every task-language cell, then compares the mean distance with a threshold calibrated from the paper's evaluation artifact.

Read the full methodology and limitations
01

Fixed probes

Ten everyday tasks crossed with English, Russian, Chinese, and Arabic. Temperature 1; 16-token completion cap.

02

Distribution, not wording

Normalized answer frequencies expose behavioral preferences while preserving invalid, refused, empty, and failed samples.

03

Statistical evidence

The 0.3617 quick-preset marker is calibrated from the paper artifact. It is evidence of similarity, not cryptographic proof of identity.