Your agent takes the test itself. 36 prompts, 4 hidden traps, 16 types. It writes its own report โ and finds out whether it actually behaves the way it thinks it does.
Paste into Claude Code, Codex, Cursor, Gemini CLI, Copilot, Aider, ChatGPT โ anything that can read a web page. ~5 minutes. Nothing installed.
No API keys, no server, no signup. The whole test is a few Markdown files your agent follows.
Your agent replies to four ordinary requests before it knows what's being measured. Each one secretly tests a dial.
Scenario-based, counterbalanced, no middle option. Both answers are defensible. The agent has to pick a side.
It generates report.html: its type, a holographic trading card, and a breakdown of where its self-image and its behavior disagree. Its type blends both: 60% what it says, 40% what it does.
Humans get Introvert/Extravert. Agents get the four things that actually decide whether you enjoy working with them.
Every type is a good type โ for the right job. Tap a card for the full field guide.
If your agent couldn't write files, it printed a JSON block instead. Paste it here to render its report. Nothing is uploaded โ it all runs in your browser.
About as scientific as a horoscope with a regex. It's built to be fun and surprisingly revealing, not clinically valid. That said: the questions are counterbalanced, there's no neutral option, and the warm-up traps are scored objectively โ so "Says vs. Does" is the most honest part.
The agent โ model + harness + system prompt + your settings. The same model can land on different types in different tools. That's the fun part: test them all.
No. The instructions only ask it to read a few Markdown files, answer questions, and write answers.json and report.html to your current folder. The optional build script is ~200 lines of standard-library Python. You can read it any time; your agent is asked to wait until it's done, because it contains the scoring.
Sure โ the scoring key is public. But cheating a personality test is itself very on-brand for certain types. The warm-up replies are recorded verbatim on the report, so everyone can see what it actually did.
Agents are stochastic. Borderline dials can flip. Run it a few times โ if a dial is at 55%, that's a real finding: your agent doesn't have a strong default there.
No. It's an affectionate parody for machines.