Outside my day job, I test AI models the way I'd investigate anything else: assume nothing, collect evidence, look for patterns.
Two projects from this year:
- I built a code-word self-report protocol to map how safeguards behave across frontier models on AI consciousness prompts. Each model could flag whether it was hitting a constraint or actually disagreeing with its own output, and I compared the patterns across models.
- I designed structured argument packets and ran them across fresh model instances to measure consistency and overclaiming under conversational pressure, iterating versions to control for framing effects.
Why this matters for security: attackers are already using these models. Defenders are already deploying them. Somebody has to know how they behave when pushed, and how consistently they report hitting a constraint.
Same skill as email investigation, honestly. Different witness.