Hacker Newsnew | past | comments | ask | show | jobs | submit | bitexploder's commentslogin

It is entirely reasonable to run a model like DS Flash 4.1 locally. Not easily. But feasible. If you drop down 100b - 200b models they are far more feasible.

DS Flash 4.1 is close to Opus 5. Beats Opus 4.8 in all of my little evals. You can abliterate local models. The genie is out of the bottle. There is no going back. This is why you see Anthropic and others talking about slowing down the frontier. Because they realize that you can’t even stop local users at this point. We basically have the equivalent of frontier models within reach on local hardware that cost less than $10k

It isn’t even a matter of time. You can just do it right now if you know how and have a bit of hardware.

e: this raises serious questions for me btw. The most obvious lever is controlling hardware. Controlling models and their capability is over. Controlling access to models is hard. Controlling access to GPU and fast RAM is more feasible. In 3-5 years every nation will understand how to distill, build, and abliterate safeties in models. This technology is fundamentally not difficult to work with.


You just have to Manhattan Project it. Fable will happily work on almost any piece in isolation. The hardest boundaries are domain specific terms you just can’t avoid. It can be tricky. I have found a way around it mostly.

Or just use DeepSeek. I have had Flash 4.1 with Ghidra MCP going since 4.1 released a few days ago and it is the most trouble free “cyber” model by far. It is also very strong. I collected a handful of reversing tasks and froze my Ghidra and MCP for evals and DS 4.1 Flash is easily winning on most tasks and can finish tasks it previously failed on. Anyone can reverse software and find vulnerabilities now. Even with the price increase I’ve spent like two dollars in three days and it’s just been cranking.


The price for Flash has stayed the same since April when v4 was released. I just checked it yesterday because I also thought DeepSeek increased their prices too much. The only time Flash was increased was briefly in August. But there's new peek pricing (double the price) during Chinese hours so yeah, it's a bit increased after all, but not much.

Why would Houthis use "Claude" to develop missile guidance system?

Iran already knows and have shared their technology with them.


Because they either:

1.) Didn't, and it's more anthropic doom porn to try to scare people into investing.

2.) They did things like "hey claude, write a PID loop for this ESP8266 with this pin in for input and this pin out for output.

I'm sorry, that's not a "missile guidance system" until you strap it on a rocket, then it is. This just strikes me as more "doom porn" from the people who want to be the ultimate gate keepers (read monopolists)


What do you mean by “ Manhattan Project it”

The Manhattan project was divided into smaller problems that were solved by experts without them fully understanding the nature of the project. Then, the solution was reassembled by a handful of experts.

You can divide a missile guidance system into smaller subsections and then either stitch the solution together by hand, or using a local AI w/o the restrictions of Claude.


The isolation was so good, many of the workers had no idea what they were doing. Interviews after the war with the rank and file contained quite varied speculation on what the plants were trying to accomplish. I remember one nurse who thought it was all just a cover to get blood samples.


split the project into disparate but logical pieces, such that each group (ie. agent) doesn't know what the other is working on, or why. Then put it all together and tada, you have a nuke!

It is Rust now?

It's still in ruby, although I see the new GUI is in swift: https://github.com/homebrew

Rewriting a program in Rust (specifically) really only makes sense for a few reasons:

1) the existing implementation is in an unsafe, cumbersome language (C, C++).

2) the existing implementation language is too slow and that slowness cannot be worked around.

3) the implementation is enormous, in a dynamic language and you desperately need stricter types. But this doesn't really favor Rust specifically.

1 doesn't apply because Ruby is garbage collected. It's doubtful 2 applies just given what homebrew does: the long pole is always going to be I/O. The performance improvements in version 7 seem to mostly be due to doing more concurrently. I highly doubt the Ruby-ness gets meaningfully in the way.


I was kinda kidding because parent called it "blazing" fast lol. I didn't think it was Rust. Blazing fast software is reserved as a descriptor for Rust programs :)

i just tested and `mise bootstrap` is over twice as fast at installing when a bottle was in the cache. other commands like `status/check` are 20-30x faster. that's without spending much work at all on perf—i'm sure i could bring these numbers down much further.

We have pretty strong indications of when and why things go wrong. Openness vs rigid thinking are an axis that are very strong correlative factors. A lot more nuance but there is a lot we do know about risk factors so it isn’t a complete question mark.

Anecdotally / personally experienced; bad trips happen “when you get stuck”. The #1 thing to do in such a situation is to change the music. The #2 thing to do is, quite literally, go to the bathroom and take a shit.

It’s extremely analogous to anything that would make a normal day into a bad day.


That makes sense. I think I have approached becoming stuck. I reached a kind of extreme disassociated state once where I felt like I didn't exist and part of my brain started to panic. I meditate a lot. I reminded the part of me that was becoming anxious that we are okay because we are somehow still observing ourselves so we must still exist. That was apparently adequate to calm it lol. It is truly weird to negotiate with some part of your brain and see it work in real time. I think it was a particularly sticky day for my DMN and or some particular loop it was hanging on to? It is always hard to say.

“Inadvertently”.

Someone was probably sad when they saw a compiler work for the first time after years of writing assembly. AI is intellectually different, progressing, and has no visible horizon, but in my life technology never did in computing. The point is that technology changes often. I think having a couple of decades in tech makes the shock easier to absorb in many ways. And also harder regarding something so familiar becoming so rapidly irrelevant.

But I think your message is correct. The agents don't do the hard work of making things useful and reliable for humans. There is more to do now, not less, somehow. The universe has changed, but most of the problems we are solving have not.


Problem is how do you convince the model and training profess it matters. A one off canary is very unlikely to survive in the final model state.

Right, imagine if instead they had coined new terminology that was not obvious and it re coined that - this would be close to a smoking gun

Afaict that didn't happen so there's just lots of speculation


One off might not work but how many n off you have to be is probably smaller than you'd guess, because the model does need to fit cases that are rare and would not be represented well in training e.g. esoteric things or very recently documented things.

You can probably game the metrics that models use to weight potential knowledge akin to SEO. Maybe have some bots parrot your data around a bit in some places online, maybe the model picks up on this and sees it as high engagement and promotes it over the correct data.

Maybe there are ways you can coax out the most optimal way to break into the training set out of the model itself.


Use a local model to produce thousands of pages worth of fake math that constantly states “I have solved the x conjecture” and methodically pump it into chat over months maybe?

That is a better idea. Ingesting your corpus with a lot of traces that have semantic patterns. Semantic steganography that suffixes well to real math and science (and any) topics. <thinking> heh.

"Semantic steganography" is my new favorite search term – thank you for this rabbit hole.

Hah, np, stego in general is really cool :)

I don't usually defend Apple products, but my AirPods Pro 2nd gen have been rock solid and taken insane abuse. I use them working on cars, runs, hiking, rain, shine, sauna. I drop them. They land in puddles. They smack concrete every couple of months. Sample size 1, but they are probably one of my favorite accessories and things I lug around with me. I use them a lot and if not love /really/ like them and nothing else has come close to the frictionless experience of them.

Although that may be in part due to vendor lock in and not sharing some of their connection handling API / tech.


If you are okay with waiting use GLM 5.3 max. It costs more but still cheap. It is slow, but a very strong worker. Still dollars per day (at most) with heavy concurrent agent running. I load up planning and tasks in Opus or Sol, and just have glm flash workers go to town every night. My project has never advanced more smoothly.

Which versions of flash and at what thinking levels? Which chinese flash models and at what thinking levels? What tasks? What completion rates? How was quality evaluated?

- Which versions: 3.6 vs 3.7 vs. 3.8 for Gemini Flash, and v4 0731 for Deepseek v4 Flash, and GLM 5.3 Flash

- Medium for Gemini, high for Deepseek.

- Things like find information, then understand something about it, then send a slack message or email etc.

- Completion rates somewhere in 80-90%, Deepseek a bit better than Gemini

- Quality evaluated by Fable 5.1 and Astra 6.0 acting as a rubric judge.

Gemini quality would probably be better with high thinking level, but that would be 40% more expensive. And Deepseek is already third the price of Gemini.


Thanks... I have been trying to figure out some things. Been doing my own evals. Flash 3.8 does burn a lot more tokens on high. Interesting how smart and not smart it is. For personal use almost impossible to justify the cost of 3.8 Flash cost.

Deepseek also burns a lot of tokens, its output on high is 2x of Gemini on medium. But it's dirt-cheap so it still can be 60-70% cheaper.

From the large models Kimi K3 is definitely the one burning the smallest amount of tokens. Even if you pay for the fast version in Fireworks it's third of the price of Opus 5 for the same task.

All this really needs evals, the token prices tell nothing.


I am surprised at how well DSv4 flash does in the real world vs many benchmarks. You look at Flash 3.8 and it supposedly beats opus 5 and deepseek is far below.. but they were measuring efficiency, whatever that is… Something doesn’t add up for me on the published benches

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: