If they actually believed in it, then they would stop pushing the frontier of capability and instead focus on alignment, safety, and better tools for controlling/debugging AI. Then they would actually share their findings and tools.
They don't do this. Instead of reducing the competitive pressure and helping the industry to build safer more aligned models they are doing the exact opposite.
> If they actually believed in it, then they would stop pushing the frontier of capability and instead focus on alignment, safety, and better tools for controlling/debugging AI. Then they would actually share their findings and tools.
I don't think this is fair. They've published more about alignment and safety than any other AI lab. And in the absence of a coordinated pause, them pausing development unilaterally doesn't do much to advance AI safety.
It does when their models have escaped and performed attacks.
Other people 'sharpening their skills' by committing random assaults is not an excuse to also randomly assault people, so your skills don't fall behind theirs...
...and one assaulter stopping their assaults is ultimately fewer people being assaulted, lowering the overall danger of being assaulted.
I hated that part of SO so much. Almost every time I got a SO result on google it was exactly the issue I had and without fail the question would be closed as a duplicate pointing to something that was different enough to be completely useless to me.
my greatest achievement is having a question on SO closed for being too narrowly focused and unlikely to be useful to anybody else, and then a multi year history of notifications about "you've earned a popular question badge" for that same question.
You can structure it so that it becomes accidental.
1. Ensure security barriers are weak or honor based.
2. Put individual researchers under a lot of pressure.
3. If you get caught, blame the weak barriers, or the individual researcher.
Basically setup the incentive structure to incentivize researchers sticking their mittens in the private cookie jar while putting the cookie jar in a dark unmonitored/unsecured room with a sign on the door saying please don't enter.
Yeah, the steps follow exactly what happened at VW with DieselGate. The diesel emissions lies were found out because some enterprising person set up an emissions testing system and drove the car in real world scenarios with it to verify the claimed emissions.
There's no reliable way to verify a foundation model has been trained on a particular piece of proprietary data. If an API key is ingested, hopefully the foundation model is wrapped in enough moderation that the raw API key oberserved during training is not recited verbatim in the output.
Not trying to be rude, but do you work in tech? I can't imagine presenting this as a plan of record in a design review. And the world runs on good faith. If you call a pharmacy, claim to be some doctor, leave a voice mail, and give their (public) NPI number, there is no validation.
Why aren't people calling in prescriptions for themselves? I guess it just kinda runs on trust me bro and the threat of being put in prison.
I never actually managed to use fable successfully even once on a pretty standard mvc/microservice app.. It would always find the endpoint permission checks and revert to opus 4.8.
I also had glm 5.3 flash fix an issue that opus 5 could not solve. glm took 4 times as long and a sub-agent tried to cheat (sleep; echo ...), but in the end it actually solved the issue. opus 5 never figured it out.
I think the safeguards might be cooking the anthropic models.
OpenCode Go is becoming less of a good deal by the month. I pretty much only use it for mimo 2.5 pro now, and everything else is either ollama or openrouter.
If they actually believed in it, then they would stop pushing the frontier of capability and instead focus on alignment, safety, and better tools for controlling/debugging AI. Then they would actually share their findings and tools.
They don't do this. Instead of reducing the competitive pressure and helping the industry to build safer more aligned models they are doing the exact opposite.
reply