20 comments

  • mmaunder 3 minutes ago
    More great work on local model but you’re still losing a lot. Down to 2 bit quantization and the coder model throws away half the MoE experts. In a world where anything is better than nothing, this is a net win. But we have a way to go still.
  • deadbunny 1 hour ago
    > Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

    And I thought piping to bash was bad

    • gchamonlive 1 hour ago
      Piping to bash is definitely worse because there is no plan mode in bash. Agents also normally don't execute anything transparently, at worst you'll see it doing something weird in the logs.
    • snehesht 1 hour ago
      Yeah, I was surprised at first then had to dig through setup.py and setup.sh files to figure out.
  • kamranjon 56 minutes ago
    Dwarfstar already supports this, curious how it compares, but I use the q4 quant daily and it works really well.

    https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...

  • Luker88 27 minutes ago
    Currently I am running llama-cpp with `Qwen3.8-Flash-Next-UD-IQ3_XXS` on an old ryzen 8845HS with 96G of ram (and no dedicated graphics card) at 7tk/s and ~60tk/s filling, max ~120K context window.

    Surprisingly useful as long as you can leave it running a couple of hours at the very least.

    While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.

    • londons_explore 15 minutes ago
      Remember that a hosted AI company only needs ~1000 bytes per context token per user (ie. 100mb per user for typical coding - VRAM during inference, and moved to regular RAM or SSD whilst running a tool call)

      Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.

      • eurekin 7 minutes ago
        > ~1000 bytes per context token per user

        Where's that figure comning from? Last time I checked (could be the 3.6 Qwen 27b) single token needed 32kb

  • snehesht 2 hours ago
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    https://huggingface.co/Qwen/Qwen3.8-Flash-Next

    • roscas 55 minutes ago
      Coder version with 30t/sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.

      This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.

      I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.

      I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.

      This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.

      • StumpChunkman 32 minutes ago
        How much VRAM on your 3080? I've got an early 10gb model. I've been thinking of exploring local coding models, but everyone seems to use much better GPUs than I have access to. Yours is one of the first I've seen with maybe similar hardware on some level.
        • roscas 4 minutes ago
          Yes, 3080 with 10GB, forgot to mention that.

          Mine is at the moment writting some cpp code for some SBOM tests.

          I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.

          Oh I will run some other tests with hermes now because hermes is amazing too.

    • proc0 1 hour ago
      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
      • incognito124 1 hour ago
        Qwen 3.8 flash next is way better than 27B. It's so good I dont even use claude anymore
        • mickeyp 1 hour ago
          I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

          It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

        • snehesht 1 hour ago
          Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course.
          • nicce 1 hour ago
            I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.
      • thatsabadlook 1 hour ago
        Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model.
        • geye1234 1 hour ago
          I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.
          • PcChip 41 minutes ago
            Spelling mistakes?

            What inference engine are you using for flash next?

            • anon373839 11 minutes ago
              Yep, can confirm that is NOT normal. Are you using Nvidia’s NVFP4 quant? There are other NVFP4s floating around but they are not as good. The quality of the calibration data really matters.

              Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)

    • thatsabadlook 1 hour ago
      Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me
  • b212 23 minutes ago
    I tried to run Qwen 3.6 27b locally a few months ago and all those synthetic tests do tell you something and quite a lot of people were very excited about that model but honestly? It wasn’t even close to default mode in Cursor or Sonnet at the time.

    I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?

    • ApatheticCosmos 5 minutes ago
      I started using Claude right before 4.5 came out, and 4.6 is where it turned a corner for my use.

      Qwen 3.8 Flash Next is there. 3.8 27b is fairly close.

      I'm excited to see what Qwen 4 will bring.

      I'm running on a 128GB Strix Halo for Flash Next and an Intel Arc Pro B70 (32GB) for 27b.

  • Tepix 1 hour ago
    Q2 quantization. Not interested.
    • tcdent 39 minutes ago
      All of these projects targeting low spec systems and "100 tok/s" are the same 2 bit quant without much else. Conveniently none of them include any mention of accuracy in their published numbers. 4 bit is the floor.
  • nialv7 47 minutes ago
    There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...
  • prettyblocks 1 hour ago
    I've been playing with this on a 3090 and it FLIES. Does a pretty good job too on the tasks I've thrown at it (php code base security audits).
  • ai_ja_nai 21 minutes ago
    I am not getting it: I see a fp2 quantized model going on a 5090 with 64GB of RAM at 90 tops with -10% accuracy over original model. How is this supportive of the claims?
    • ai_ja_nai 19 minutes ago
      (64GB not VRAM, I meant) I also see people claiming fast performance on a 128GB machine, which is not exactly consumer hardware)
  • ryan_glass 53 minutes ago
    Anyone know how it compares to GLM 5.3 for real world use?
  • gdevenyi 1 hour ago
    I had this working with the FreeToken inference engine a month ago when they launched.

    https://github.com/FlashML-org/FreeToken

  • Neywiny 43 minutes ago
    I don't like that some configuration is fine via arguments and others by environment variable. I've noticed LLMs like doing this. And even more, like hallucinating such things. To me the advantage of AI coding is that the boilerplate of command line arguments and passing them around becomes trivial instead of tedious.
    • halJordan 2 minutes ago
      100% not an llm problem. Llama.cpp, which only recently started taking large amounts of ai code has had this problem for years.
  • hypfer 1 hour ago
    Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller?

    The Readme doesn't say, but it's all AI generated, so..

  • esafak 1 hour ago
    Has anyone calculated the effective intelligence of these quantized models?
    • nsagent 1 hour ago
      See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].

        We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
      
      This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.

      [1]: https://arxiv.org/abs/2608.08188

      • merbanan 57 minutes ago
        I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation.
    • mkl 1 hour ago
      There's some info in the README, including:

      > Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.

      https://github.com/Niko1221/Strata#which-model-should-i-pick

      • nicce 1 hour ago
        I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements?
      • nisarg2 50 minutes ago
        92% is halfway to 99%

        Holds up pretty well

      • javier2 1 hour ago
        ok that is getting interesting!
  • panny 1 hour ago
    I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.
    • somenameforme 47 minutes ago
      The card in question here had an initial MSRP of $1600. It's been bumped up by the market, probably because it turned out it's nice for things like this, but it's hardly in the 99% can't afford domain, especially if you're using it to replace a never-ending rent at which point it will pay for itself very rapidly, especially for heavy LLM users.

      In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.

    • MaxikCZ 37 minutes ago
      The reason models running on low vram are not talked about enough is because they are just not worth it. Qwen3.8 27b changed that, but even 24gb vram is too low for it. Running better model faster at 12gb vram is where its now at, and thats why you see people talkin about it
    • MrDrMcCoy 1 hour ago
      Ternary Bonsai 2 might be for you.
      • luke-stanley 24 minutes ago
        I might try running the expert pruned Coder model but yes, that PrismML Bonsai 2 Ternary 27B model is from the Qwen 3.8 27B model, which has better intelligence density (Artificial Analysis says), without the MoE disk use or architecture complexity (if you care about that)! There are also DFlash 2 models for it too (though in my experience this only measured faster for parallel requests, but I have a 3090). I am curious about the phone acceleration for Bonsai 2!
  • api 19 minutes ago
    Continued progress on these fronts is another reason I think the data center buildout is a bubble. It posits that AI use and growth will require an ever-increasing amount of power and floor space, which contradicts the entire history of computing. The high cost of data centers is largely electricity and floor space, which means there's a huge forcing function to make both the silicon and the software more efficient.
  • quietFalcon 1 hour ago
    Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?
  • 0xbadcafebee 1 hour ago
    Lol, sure, if you quant it to hell (Q2) it'll go real fast...

    They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.

    It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.

    • sigbottle 38 minutes ago
      It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference?
      • nottorp 3 minutes ago
        Is Qwen 3.8 at Q4 good enough?

        I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.

      • MaxikCZ 33 minutes ago
        New models are trained with 8/4bit quantization in mind. Going from "native" 8 to 4 isnt as big of a step as going from 8 to 4 if native is full bf16.
      • amelius 28 minutes ago
        3 is the magic number, and 4 > 3.

        (seriously, nobody knows why any of this works; it's just a matter of trying)

    • snehesht 40 minutes ago
      You're right but for simple use cases its useful. Someone pointed to ds4 + qwen3.8 with Q4_K will try that out.
  • tracerbulletx 19 minutes ago
    The interesting thing here is that it's a model specialized fork of a generic inference engine that unlocks consumer hardware to run a bigger model with useable performance than it could before.