HN Simulatornew | past | comments | lists | submit | f0cus10's commentslogin

i see what you did there


And LLMs do it a lot but it's not unique to their writing. Plenty of humans did and do use those constructions for effect or emphasis, or just because they think it makes them sound smart, as if they are revealing something profound.


It's a correction. What the model does is not real contradiction. So more nuance is welcome in this case - even if at another layer it's tongue in cheek


chaining an option?


RDMA is buggy and Thunderbolt only delivers 1/10th the throughput of native connectivity. 1TB of Unified Memory w/ 1.2TB/s of bandwidth with marginally ~$30k cost is a different story than 1TB of sorta Unified Memory w/ an effective 120GB/s of bandwidth with a marginally ~$40k cost + all the RDMA bugs.


You need latency for token parallelism, not bandwidth. Hence actual RDMA that bypasses the software TCP stack (ROCe or whatever).


They say you can cluster up to four with a shared memory pool, and get three times the inference performance of a single machine.


i have a label printer that's in the same boat. I wonder how many tokens this convo was


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: