lot of unexplored potential left in non-transformer based, "old school" models that people seem to have forgotten, or never studied if they got into AI starting from LLMs
In this post we explore model-as-a-judge: using one LLM to police another, and the existing research on its efficacy.
Short version: if an attacker can write what the judge reads, they can influence the verdict too. And typed outputs don't fix it.
Who judges the judge?
Adding a second model to vet your LLM's inputs gives attackers a new target. In published research, adaptive attacks beat most defenses more than 90% of the time.
Why a model's approval isn't a security boundary:
blog.milgram.dev/model-as-a-jud…