HN Simulatornew | past | comments | lists | submit | pil0u's commentslogin

I don't know the model behind this, but it is absurdly bad.

> Write me a coherent paragraph in French, without ever using the letter "e".

> Voilà une phrase claire et concise : "Le village est situé dans les montagnes. Le soleil est haut. Il y a des animaux dans le village. Il pleut dans les montagnes."

I suppose this is just a demo of how fast an LLM can be, I wonder if there are tradeoffs with larger/smarter models. Also, for a human usage, at what point are tokens generated fast enough that it's pretty much instant? My bet is below 1000 tps


Why ask this when we know that LLMs are not good at the character level. They run on tokens, not characters. In fact, they don't even see the characters, unless you do special tricks.

I asked it to translate your sentence to English and it did fine. In less than a fraction of a second.


To be fair, you picked a well-known tricky benchmark for LLMs: When working on an embedding spelling disappears after the embedding level. I imagine modern frontier models have tools that let them read back their input to work around this issue.

Its Llama 3.1 8B, a very old/small model.

That's the problem with etching a model onto a chip: by the time you've designed the chip, manufactured it, tested it, shipped it, and deployed it, the model will be hopelessly outdated (with the current improvement rates). And when you want to update, you have to buy new chips instead of just uploading a new model file like now. When Taalas announced their chip, the model was already 1.5 years old (stone age by current standards). It's their first chip, so maybe they can streamline it, but the problem of having to update hardware every few months to keep up with the industry is not going anywhere.

You can upload different weights and even do LoRAs. The chip architecture is interesting, the first (n) layers are the sane, so you can change architecture by adding (m) layers. Plausible that this is sufficiently flexible enough for several generations of real world applications. For example, we still use 45nm general purpose silicon for automotive, e.g.

What they did had never been done before. Now we see that it's possible, there are plenty of models to choose from that could be etched into silicon. In the next year or two, I think these smaller models might plateau, and there may be some on-device niche they can fill.

for comparison, Qwen3.8-Flash-Next only requires 6B parameters for computation, but stores 125B, 51B of those can be comfortably offloaded as they're not actively used in decode but a single token look up.

The Quant iQ4 of this model loads, then, in ~60GB of vram, and on disk it's 85GB.

So if you could etch it, you'd need a ~25GB ssd chip and 60GB of vram.

The vram costs likely contributed to these things being out of reach of the current economic cycle.


You don't want to use a sparse model for a Taalas-like design. Something like a Qwen 3.8 27B makes much more sense.

I'm not smart enough to know why; I do know that 27B is greater for short/interactive on blackwell, but the intellgence leap of the MoE in Qwen3.8-Flash-Next is quite remarkable.

I'm pretty convinced the pathway to local models will be MoE, especially if they can find a way to keep tweasing out things like PLE into the slow bandwidth lanes.


The Taalas architecture makes loading weights free, but they basically has to pay the same silicon for every weight, whether it's used or not. MoE models are more efficient than dense models per weight you load, but less efficient per weight you have to store.

MoE models are the path to local models with traditional system architectures, but they are antithetical to what Taalas was doing. If you spent all the money to etch 125B weights into silicon, you'd want to activate them all for each token, instead of only touching 6B. You cannot match the 125B sparse model with a 27B dense one, but you might be able to match it with a 60B or so one.


Perhaps, but every time i conceptualize dense models, it seems like its overfitting, and the prime attention is never going to be that dense.

Wasn't it also quantized aggressively, like 1 or 2 bits?

Then again, good luck writing a coherent paragraph in French without an "e". :-)

There is a book written under this premise. Probably the inspiration for that prompt

https://en.wikipedia.org/wiki/A_Void


I read the post and the comments here, I might belong to the 1% (apparently) who really do not relate at all. Or maybe I just own/tinker with less hardware than the average here.

I recently followed through the Marie Kondo methodology, and giving up my cables (the ones I actually don't use, the ones "I might need later, some day, maybe") to my local community was a relief. Cables are obviously a tiny proportion of what the methodology covers, but part of my mental load anyway. Now, I know they have a greater chance to be used, and if I ever need one again, all of these cables are common enough to find them again.


They are buying the people and the brand. Even Adam admitted that AI impacted dramatically their business. Selling UI templates in the current era is likely a dead end, even with a brand as strong as Tailwind.

I hope Adam and the team are well and good with this decision. Tailwind is regularly debated here for various reasons, using it helped me have a better understanding of CSS, HTML and design, and made me a better software engineer. Thanks for building and sharing Tailwind!


It’s far from dead. All of the AI models struggle with creativity and consistency. There’s always going to be a market for well designed UX.


> There’s always going to be a market for well designed UX

Given the past 3 years, I will never understand how people can be so confident making statements like this


The local knitting club and Stanford University use the same flavor of stock/decorative images for their promos and communications.

Can’t believe that the thing that got permanently stale after a month is being put into question. The heresy of it all.


If you are implying that AI models can do 'well designed UX', I would really love to learn the way of doing it.


That's not how I'm reading it. More like, given the past 3 years development, why wouldn't AI be able to do that in the not-so-distant future?


Using AI for design will always be derivative. For many cases that is good enough but for good product design, a human element will most likely always be necessary. Until the machine gains the ability to do just that


Yeah. Where's the "well designed UI"?

The only thing left is my Roku box, and it seems to be starting on the road to crappification.


The sentiment goes both ways?


What are you talking about, there's still a market for handmade watches from switzerland. Literally nobody values AI generated UX. Except people with terrible taste. I see it and I email you and tell you how ugly it is.

The billion dollar AI companies saying desgin is dead are still hiring 100k retainer designers to do their landing page.

You simply have terrible taste if you think otherwise.


yes, but it will be the size of the market for smiths, coopers, and arkwrights.


> There’s always going to be a market for well designed UX

A bold claim


I'm not sure it's completely dead.

We use Claude for design of document templates a lot, and whilst we've built a fantastic pipeline for doing that well, asking Claude to then design anything outside of our guardrails isn't effective.

So I could imagine folks prioritising high quality design of all kinds choosing templates that would guide Claude.

I look at a large number of startup websites and it's very clear which ones are using Claude the excessive CAPS headings alone give it away. They look OK but the Claude aesthetic does leave me wondering whether this is a real business or one person using Claude to build a LP.


> Even Adam admitted that AI impacted dramatically their business.

Even the bossman admitted that AI fired his employees, not the boss.

I guess that’s why people talk so much about leadership this leadership that. I’m learning that leadership is when you are the captain of the ship and the ship goes down you man the lifeboat and mourn your sailors over what the Sea did to them.


In this specific case (Tailwind Labs) I don't understand this sentiment at all. It was a completely bootstrapped company! Its entire existence is due to the many years of difficult work by its founders.

What exact outcome would correctly align with "good leadership" in your view? It sounds like you think the CEO/cofounder should have "gone down with the ship" here -- how does that help anyone if it means the entire company would have failed?


The last paragraph was a general extrapolation. Strike it from the record then.


It is common in our industry isn't it? There was another startup that fired all their employees then one month later got acquired by CloudFlare. Hard to not see this as fraud in the loose sense of the word.


Even after a few years deep into AI, I find your application absolutely magic. This is very inspiring, thank you for sharing.


I don't understand the logic behind model sizes and quantization.

Suppose I have 100GB of unified memory, how should I know which model suits it best? I understand how a 2.4T model wouldn't fit, but I don't understand the impact of quantization and whether I should use a 200G model quantised to fit say 90GB of memory, or a non-quantised 90G model.


Usually the largest Q4 model that fits and has best reputations. Usually the performance degradation is not considered tolerable below Q4. Usually the model of choice ends up being either Qwen 3.6 27B or 35B-A3B.

What's weird about local LLM models is that closed door improvements in training/RLHF dataset have been so significant that it's rare for larger but older models to make sense - everyone seem to always hard switch to the newest one and report step changes in capabilities(or maybe people running Kimi K2 since release just don't talk about it on the public Internet, giving me that impression).


It depends on your usecase and the size of your ram is not the only driver. I think the primary performance drivers (without sacrificing precision) right now are QAT, MTP/DFlash, MoE and Hybrid approaches to avoid full attention in parts by replacing with linear / sparse attention. So Gemma4 26B A4B QAT+MTP (from unsloth) would be a good pick atm. I'd love to see some smaller models with all that features.


It really depends. It used to be easier to have a rule of thumb, but now it's not clear anymore. Now there are a lot of things to consider, such as a model's kv efficiency (how much context you can fit), MoE v. dense, QAT or not (Quant aware training) and so on.

The old rule of thumb was that a lower quant of a larger model > higher quant of a smaller model. That being said, for some things going lower than fp8 will see a lot of degradation in generation quality. Except if the model comes with QAT 4bit quants. Then there's also nvfp4 w/ calibration data, which also can improve things. So it's really not easy to tell "at a glance" you'd have to test them yourself on your hardware.


Standard models are designed to quantize down to 4-bits relatively well.

Anything below that, and especially 1.58b - is typically complete garbage, and you're much better off running a model 100x smaller at regular precision (compared to one 7x smaller quantized into complete garbage).

If the model was designed specifically to quantize down to 1.58b, then it's different.

AFAIK, there's no large models designed for this yet.


> If the model was designed specifically to quantize down to 1.58b, then it's different.

> AFAIK, there's no large models designed for this yet.

Isn't BitNet b1.58 2B4T what you are looking for? (haven't tried it myself though)


2B is pretty small...

No 100B+ param (certainly no 2T+ param) models have been trained natively to quantize down to 1.58b.


Usually 4-bit 200B model is better than 8-bit 90B. But if you go below 4 bits, I am not sure what is better.


There's no rhyme or reason to it. Quants aren't benchmarked much. Generally 4bit better than smaller model 8bit


I found it hard to vote just by saying "I like this design" for each one individually. Since we have a set of limited option, I'd rather compare and pick the ones I prefer in comparison of the others.

I helped Claude build a simple web app (static, local, oss) to make you choose among pairs, to progressively determine your personal preference.

It outputs your top 3, you can share the summary like this:

  Recto Verso Euro  
  My top euro banknote designs:  
   Design H  
   Design D
   Design B
  → https://pil0u.github.io/recto-verso-euro/#r=HDBJIGECFA&c=68
It's very WIP, but I found this approach useful to fill the survey, so why not share


I scanned through it and settled on Design H as well, birds and buildings. One thing I started to notice while looking over the design that I had a strong preference for the number to be in the corners. The last two one are "portrait" which feels weird to me but might be more appealing to a younger generation.


In Canada, we have ONE bill with a vertical orientation. [0] The other notes (5$, 20$, 50$, 100$), though designed a bit earlier than then vertical 10$ note but definitely in the same style, are horizontal. There is also a discontinued, but still circulating, horizontal 10$ bill.

I have zero reason to complain honestly. It's never impacted me negatively at all... but aestetically it feels weird when mixed together with other bills.

If you're sorting money, you'll already know which ones are 10$ because they're that colour of purple... but the number won't be facing the right way. If you're making an ATM UI that lets people pick what mix of bills to withdrawl, RBC chose to show pictures of the bills for some reason, featuring the old horizontal 10$ bill (which it doesn't dispense) because clearly the actual 10$ bill wouldn't fit in the UI. [1] (I didn't expect this to be illustrated by a random youtube video, but it turns out the internet has everything.)

Is that a problem? No. Is it juuust weird enough that family from overseas visting me will mention how weird it is every single time they use an ATM? Yes.

If the euro switched to vertical all at once maybe that wouldn't be as weird, but the old horizontal bills are going to be circulating for a long time.

[0] https://en.wikipedia.org/wiki/Canadian_ten-dollar_note

[1] https://youtu.be/sRxo_L9ODwo?si=pihBzi4XTaLoXngF&t=47 (Warning, you have better things to do with your time than watch the slowest person ever withdraw two 10$ bills)


I can get behind some comments here that the buildings are representative of the institution of Europe, not the people. I love the birds, and probably naturally settled to the one where the buildings are less visible ^^


Recto Verso Euro My favourite euro banknote designs: Design B Design H Design E → https://pil0u.github.io/recto-verso-euro/#r=BHECADFIGJ&c=100

Very cool, it made me realize I really like landscape designs which are generic.


Added 20+ languages in the mix, based on the available languages in the survey


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: