> They were building in stealth for 2 years, I was building in stealth for 2 hours…
> Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe.
flockonus 4 hours ago [-]
No question OSS is amazing, but this video is a satire at best.
It doesn't take much attention to see the results on right vs. left side are significantly different.
Jev is not interesting if it's not "smart", a 1B param model is most definitely not smart.
nullbio 22 minutes ago [-]
Jev has a 32k context window. I doubt it's a large model.
10/10 programming language detection
9/10 human language detection
10/12 unit magnitude comparison
All incorrect answers are marked with low-P.
It (DiffusionGemma with the Jev mode) can also solve an ASCII maze.
vrc 1 hours ago [-]
Out of curiosity and semi unrelated — why do so many of these projects with customized encoder-decoder setups use earlier Qwen versions like 2.5 and 3 and not the smallest 3.5? Purely the few 100m params, or something else in the latter’s arch or pretraining?
augment_me 48 minutes ago [-]
In my experience if you tell Claude to port LLM-like stuff without explicit steering for versioning, it will default to the most popular thing for this in its training window to reduce errors. 3.5 is outside its training data.
FuckButtons 38 minutes ago [-]
My guess would be qwen 2.5 predates linear attention which would be more complex to use.
> They were building in stealth for 2 years, I was building in stealth for 2 hours…
> Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe.
Jev is not interesting if it's not "smart", a 1B param model is most definitely not smart.
Runs ~0.2s per decision on my DGX Spark.
All incorrect answers are marked with low-P.It (DiffusionGemma with the Jev mode) can also solve an ASCII maze.