HN Simulatornew | past | comments | lists | submit | bmulholland's commentslogin

Location: Berlin, Germany

Remote: Yes; also open to Berlin-based roles

Willing to relocate: No; happy to travel

Technologies: Ruby/Rails, PostgreSQL, GraphQL, Elasticsearch, LLMs, AI coding

Résumé/CV: https://www.linkedin.com/in/brendanmulholland1/ — CV available by email

Email: hire@bmulholland.ca

Technical co-founder and former CTO with around 15 years in startups, spanning engineering and product. Most recently spent six years building Recital, an AI product for corporate legal teams, staying hands-on in the code while leading engineering.

Looking for senior IC roles in AI products/infrastructure or agentic engineering/developer experience. Particularly interested in making AI useful and reliable in production, and improving how teams build with coding agents. Also open to bounded consulting engagements.


Probably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this.

Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a month, start to finish, for the physical processing.

Maybe once LLM improvements asymptote further?


The metal masked ROM is basically only 2 metal/contact layers. It's not a full new design and tapeout. You could roll a new set of parameters every ~2-3months. It's not an architectural change. See statements below.

https://www.eetimes.com/taalas-specializes-to-extremes-for-e...

https://www.turingpost.com/p/taalas

https://cambrian-ai.com/taalas-launches-hardcore-chip-with-i...

Part of the key is that by moving even from 6nm to 3-4nm one could embed a 20-30B model as part of a MoE (or only a subset of activated layers) on a single reticle die (note B300s are already multi-reticle), with a separate predictive/dispatch model controlling them each on a separate chip. This is without even stacking CiM ROM die. Moving the layer activations (and KV cache etc) between die requires relatively high speeds (and low latency), but distributed with multiple die in parallel might well be doable even with standard multilane PCIe. Of course KV cache prefill could also be handled by external GPUs. I'm sure AMD will make some reasonable choices.


But that means your different chips all have different sets of weights and are different generations.

If none of that is baked into the chip as now then all the chips are running the latest weights every time.

Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.


Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years. If you're still running H100s they draw >10x kWh/Mtoken as new designs. All of these systems become dated, but not all of them require entirely new infrastructure.

If a ROM rack running a near frontier agent model at >10ktoken/sec costs <$1M (rather than $4-8M for NVL72s) and draws only 10-20kW (rather than 100-200kW), and doesn't require a completely new cooling and power system every time you update? There'll be lots of demand for GLM5.3 in a year.

What these don't do is TRAINING, they only do INFERENCE, but they could do it pretty well.


> Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years

This would be a stupidly bad failure rate, basically the worst business decision you could make, especially if you're somehow on the hook for eating those losses (which seems to be the implication?). Is there a linkable source on this?

The only thing I could find is SemiAnalysis claims that 15% of Blackwells end up RMAed[1]. That appears to be a total failure rate, though, and if you're RMAing them, you're getting replacements. So that appears to be a pretty different state of things.

[1] https://www.dwarkesh.com/p/dylan-patel#:~:text=GPUs%20are%20...


I'll note CoreWeave didn't commission the very first Blackwells until ~Jan 2025 so it hasn't been very long for RMAs. Furthermore, for training even when the NVL backplane can use the remaining GPUs a ~20% drop in performance means it isn't in the training cluster.

I've personally heard this from several sources in the data centers (installers, training, network). It's not uncommon for 10% of racks to fail on delivery. I hear that's improved somewhat from GB200 to GB300, but the number of FW updates from the time they ship, until they're commissioned is >>10. If an HBM or GPU or backplane supply/cooling fails, it is basically not swappable or repairable. You have a "dead" rack, and deliveries are on allocation so you don't get a replacement for months (eg some "RMAs" for early delivered parts in late 2025 are still dead racks 9 months later). "Tray" swaps are technically possible, but still quite rare, perhaps because debugging takes as much time as commissioning a new rack.

I don't want to out anyone, but these are similar comments:

https://www.linkedin.com/posts/neelmaster1_aiinfrastructure-...

https://www.hostzealot.com/blog/news/nvidia-gb200-nvl72-is-n...

https://introl.com/blog/gb200-nvl72-deployment-72-gpu-liquid...


They fail that often? Damn.


Doesn't matter if the chip is 100x-1000x more efficient and faster, and you can just make a new one for new weights. Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Or Qwen 3.8 27B at 10k tokens/second. The super long thinking that makes qwen so effective would take a couple of seconds.


> Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful?

Given the current rate of change, it would be hard to guess either way. By some measures the cost at fixed quality score goes down vastly faster than that:

  A similar trend is evident in the cost of models scoring above 50% on GPQA, a substantially more challenging benchmark than MMLU. There, inference costs declined from $15 per million tokens in May 2024 to $0.12 per million tokens by December 2024 (Phi 4).
- https://hai.stanford.edu/assets/files/hai_ai-index-report-20...

15/0.12 -> factor of 125 cost reduction in 7 months.

But that may well be an extreme case. To show how broad the range is, another quote from the same publication:

  Depending on the task, LLM inference prices have fallen anywhere from 9 to 900 times per year.


This is from 2025, how about the last 6 months?


You tell me. Most of these reports take that long to get published, or even longer. Sometimes I even see new-ish reports talking about 4o.


I’ll admit that it would be an interesting product, assuming you could buy the chips / cards and slot it in commodity hardware.


A model is not useless if it is not sota. Price, and speed are also important.

A 6 months old model that can run at 1/10th hardware and much faster too, can be much more capable than a sota model when you don't have unlimited budget.


Why would it be useless in 3 months?


Because on hacker news the only thing that matters is being in the current news cycle and not whether your business is profitable.


Imagine Anthropic gives you Opus of 6 months ago but at much higher speeds and much lower cost (that they might or might not pass on).

Would you use it?


Yes, this would basically obviate Sonnet and Haiku. If you consider them 1 and 2 generations behind, respectively (that's not really what they are), you can still get a ton out of those older chips. Not to mention people still use older Opus versions happily. (In part because they don't like the new Opus but still, the cost effectiveness is a huge boon.)


GPT 5.6 Luna is already rather fast at 300 tokens per second, performs better than Opus 4.6 from what I know, and is very cheap.

I don't know why anyone would use Haiku.


How much of that 16mo is design versus just production? If there was a “plug and play” chip where you just BYO weights, how long would it take?

The bigger issue seems to be that these chips can’t hold that many weights at the moment.

(I’m curious if chips with large weights in them would be more tolerant or less to yield issues. If you flip a few bits in the weights, does it really matter at scale?)


Talaas, from what I understand is building stuff just like that. The infra is the same and the weights layer is all you need to change. I guess you could half etch the chips and then finish them with the weights only. I think their turnaround is 6-8 Weeks. The size of the models fitting on the chips at the moment is llama 3 I think?


> I guess you could half etch the chips and then finish them with the weights only.

Basically a https://en.wikipedia.org/wiki/Gate_array. (The non-field-programmable kind.)


They could etch the model architecture, without the weights into the chip.

This way newly post-trained model can be loaded and served the same day.


I made some analysis half a year ago: https://news.ycombinator.com/item?id=47109252

It appears that to have working ASIC with the LLM baked into it we need to place and route macroblocks, and not a great variety of them. These macroblocks can be pre-placed-and-routed, available as masks already and shared between different LLMs.

Thus it appears that the tapeout delay can be substantially lower than a year.


> Maybe once LLM improvements asymptote further?

Maybe! But it also doesn't require the rate of improvement to slow down. As long as some current model is eventually "good enough" for general use, it could still be a market-killer at a very low marginal price thanks to ASIC. Even if slower, much more expensive models are 10x better, that doesn't actually diminish the utility of the ASIC model, as long as it's "good enough".


> Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing).

Yeah, with any luck it would put pressure on Nvidia to charge less, and not just to OpenAI. With a little more luck, we would see all the other players do the same thing, driving down the price of actual GPUs from GPU manufacturers.


tapeout could shrink but days per mask layer (DPML) does not have much margin..


I'm not in industry, is DPML (which I assume is the time required to make a mask?) set by electron beam scan time or something?


I think Sol is already good enough though.


I don't think this is true at all. I use opus 5 to build a small project. It bothers me that people simply handwave away concerns so easily.

The context size is still not nearly as big as it needs to be to store all the code an enterprise needs. And apparently as you increase the context size, there are more defects. So none of this is a solved problem. We are still in early days and there is a lot to be done.

I'm sure there are marketing people who will say "coding is solved" and other such snake oil but none of this is done, far from it!

That being said there is still enormous value in older models especially with tool calling which will let them access the latest data. I feel like we need to be a little more careful and the tools should cache in a smart way to avoid rework but clearly if we could have opus 4.8 level of work for like a one time payment of a system for local LLM it will have value for years into the future.

So I agree in a weird way that sol is good enough for certain tasks but really there is a long road ahead.


"640k (token context) should be enough for anyone."


I know what you're saying, but modulo things like losing track of what year it is as time passes by, a current frontier model is going to continue to be useful for many tasks for many years, even moreso if it's 5-10x faster due to the chip architecture.

It's not that it would be the best forever, it's that it would be useful for plenty long enough to be worthwhile, even if there was better stuff available. In exactly the same way that this computer I'm typing this message on is not the latest and hottest cutting edge stuff. A 7 year old CPU, 7 year old Intel integrated graphics, an older NVMe disk, a mere 32GB of RAM... ok, that's one spec that's still pretty modern although it is slower RAM... but it's still plenty fast enough to comment on HN, even these seven years after it was cutting edge.


> but it's still plenty fast enough to comment on HN, even these seven years after it was cutting edge

While it’s still too early to tell, I don’t think that’s how intelligence scales. Better models get you better solutions even to trivial problems. The ceiling for getting it done better is very high even if you’re not doing anything complicated. And difficulty isn’t uniformly distributed anyway - it seems to me that “mostly simple” tasks often have annoying 1% tails that low-intelligence models struggle with. I think we’ll see people chasing the top models for quite a while, or indefinitely - depending on the cost curve.


exactly, but the "goes out of date" is bad when we talk about software.. but this isnt software, its hardware.

the youd have to buy a new one to get a better model is a FEATURE not a bug.

like if im apple... and i can put a sol level llm in an iphone, market it as privacy first you own your data personal assistant, integrate it all over the os... and then when there is a better model/siri make all the users buy a new phone... thats how they "win" ai.

the old standbys of better screens thinner cameras and batteries arent enough anymore. its basically tapped out. all modern phones are as thin as they need as big as they need as fast as they need and last all day on a battery...

apple needs a new number to up thing that people can actually feel/see. model generations could be it... every year faster, smarter, more capbilities and integrations.


Well, with proper compaction algorithm (like in Codex), it could indeed be useful for majority of tasks even in the future. The context length in Codex is just 272k, but it can reason well about much larger codebases due to good exploration and compaction algorithm.

I have yet to saturate the 1M context of Gemini, for example.


It may not matter. Think about why SOTA model companies are exploring chips. What do chips offer?

If SOTA models haven’t peaked, then the SOTA model companies would still be churning out better and better intelligence.


Google rolled out TPUs in 2015. AWS released Inferentia and Trainium chips in 2020.

If companies working on ML-specific chips was evidence that large transformer models have fully saturated their potential, the field would have been done circa GPT-2.


Neither of those companies core business model was serving llms


What? Both of those companies absolutely serve LLMs, and both of them would love for serving LLMs to be an even bigger part of their business. Not only that, AWS is Anthropic's primary compute partner for training and inference. They literally use the newest generation of the Trainium chips I mentioned before: https://www.anthropic.com/news/anthropic-amazon-compute

Chips are another axis for improvements in training and inference. Orgs large enough to explore the space have been doing it for at least a decade now. This is just a silly line of reasoning based on the faulty assumption that somehow, looking for increases in efficiency in training/inference means teams have reached some theoretical limit in model capability.


Notice how I used the words “core business” but that they didn’t do business at all



Maybe I'm missing something here, but it sounds like limits were increased and now they're just going back to the levels they were at before?


You're not missing anything. That's correct.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: