Adam Holter Tests: My AI Model Testing Archive
Browse Adam Holter's AI model tests, prompts, screenshots, interactive HTML results, and model-specific archives.
Browse Adam Holter's AI model tests, prompts, screenshots, interactive HTML results, and model-specific archives.
Browse Adam Holter's AI model tests, prompts, screenshots, interactive HTML results, and model-specific archives.
Thesis is Every's existing editorial column, not a newly launched subscription.
How AI coding-agent permission modes trade off safety, human review, and uninterrupted work.
AI model rankings, benchmark results, cost-efficiency comparisons, and Codex growth charts.
A problem I've noticed with even frontier models that I don't see discussed very often is prompt leakage.
With the rise of token maxing, enterprise companies are getting pretty big bills to pay and running into...
Composer 2.5 is exactly why open source models can never win. People often look at random benchmarks, compare...
AI will have negative effects on our society and I am not afraid to say it. It already...
Matt Walsh is wrong about AI. First off, you have no right to employment. Rights are inherent to...
OpenAI’s general purpose reasoning model has disproved a central conjecture on the planar unit distance problem. The model...
“Google is a leaky bucket.” Just like their Pixel phones always leak, there have been tons of leaks...
Pay attention to Google’s latest Threat Intelligence Group report. It documents what appears to be the first case...
You shouldn’t use GPT-5.5 Instant. The reason is straightforward. You should not use any instant models if you...
SubQ claims the first fully subquadratic LLM tuned for 12 million token reasoning without hybrids or quality drop....
I should have applied to the GPT-5.5 party anyway. The RSVP form went live and I passed because...
Andon Labs released Blueprint-Bench 2. The results show the first measurable signs of 3D spatial intelligence in frontier...
This GPT Image 2 prompt is going insanely viral right now. The full text tells the model to...
OpenAI traced their models increasing use of goblin and gremlin metaphors back to rewards given during training for...
I pay for both the twenty dollar OpenAI plan and the one hundred dollar Claude plan right now....
OpenAI told GPT-5.5 twice not to talk about goblins. The Codex models.json file contains the exact same instruction...
Anthropic launched Claude Design on April 17, 2026. The tool lets users create prototypes, slides, one-pagers and branded...
Opus 4.7 correctly identified Kelsey Piper as the author of 1000 words from an unpublished heist novel. The...
GPT-5.4 can play 5D chess. It beat Claude 4.6 Opus in a match that ran close to five...
Sayash Kapoor and team just released results from their first CRUX evaluation. An AI agent built a basic...
I’m an addict. A 1% addict. When a new model releases and a benchmark score moves from 90...
OpenAI launched the Codex super app on April 16 2026. The desktop update merges ChatGPT integration through background...
Claude Opus 4.7 launched on April 16 2026. It improves on Opus 4.6 in software engineering with particular...
Claude announced the Code desktop redesign on April 14 2026. The main change centers on a new sidebar...