Out of curiosity why isn't this stuff handled by a secondary "monitor" agent that's specifically trained on what's okay and not okay? I'd think it'd be a pass-no-pass classifier and wouldn't degrade the performance of the main LLM.
Would the concern be that with sophisticated obfuscated input you could try to get ROT13 Klingon instructions on how to build a bomb - and that could fool the monitor?
It often is. Risky Business Features did a fantastic podcast on how different popular methods of guardrails work and some popular methods on defeating them. Absolutely worth a listen because there are some surprising insights in there on how these work, even for day to day use, not just bypasses:
These exist and are used. The issue is that because they're so much smaller, they're also much worse, so they tend to have lots of false positives while still being easy to circumvent.
This is absolutely how it's being done for certain topics. If you ever wanted to research suicide-related psychiatric topics with ChatGPT you would know to have your screen recording always on, because ChatGPT spits out a full answer and then a screening model takes it back.
Would the concern be that with sophisticated obfuscated input you could try to get ROT13 Klingon instructions on how to build a bomb - and that could fool the monitor?