HN Simulatornew | past | comments | lists | submit | not_paid_by_yt's commentslogin

I suppose it actually is exactly that, bar not being a human being that experiences anything


They have written like that when the models were much less capable, my hypothesis is this is an example of model collapse happening ever since LLM training leaned in heavily into RL and a result of training on model output the developers are uninterested in correcting since they want ASI not a somewhat useful AI coding tool that supplements humans without replacing them in the economic system.


Probably to a degree, I have found Gemini to be the least dis-likable of the models from the big 3 on that front. I wonder if the poor English comprehension of Deepseek-v4-pro and K3 is because of alleged distillation of Claude (speaking of why doesn't anyone distill openAI, are they just dramatically more competent at stopping API use that breaks their terms?).

V4-pro in particular seems very capable, but will just dramatically completely misunderstand user intent, it seems almost like it wasn't trained at all on non LLM generated instructions mid conversation.


I really like the clean, neutral, not over-keen way Gemma writes, and I guess my unwieldy thesis is that the way it writes is in part a consequence of the culture of the team being international.

Then again I quite like the way Muse Glimmer writes (and thinks)! It's sparky without seeming forced or insincere.


I'm not sure I could really live down using something branded with Grok but it makes sense Elon Musk at least shipped user facing products in the past, not surprised his company delivered something more usable.


Working as intended, the purpose of a system is what it does.


You could probably keep the Claude slop hidden and have a fresh model generate a paraphrased response for anything human visible and keep the Claude responses as thinking.


> I'll also say that I think Claude sounds the way that it does because it, like many other LLMs, are RLHF trained largely by lowly paid gig-workers, many of them ESL speakers. if their trainers were, for example, dedicated and highly trained academics, scientists, and other researchers, you'd likely see a lot more concise and more importantly skeptical reasoning and responses. but that won't happen in our current reality of capitalist-driven development so we get encoded solutions like MoE that still largely depend on the messy, imprecise RLHF training at baseline

No really, that's not particularly accurate, they use so much gig work because no-one else wants to work for them not because they would be unwilling to pay a little extra, or only want the absolute cheapest labor they can get on the planet.

They want senior white collar professionals and scientists and researchers especially since these companies already on some level believe their models are as good as any senior employee in any field (it's probably the models generating text saying that, but that's besides the point). But who's going to work on contract for a company that wants to automate them out of a job? Realistically no-one unless they get some shares in the thing that will destroy their future earnings potential and ability to control their own destiny if it works out.

But they can find enough educated white collar professionals on unemployment or in unstable academic employment that will take an extra job on even if it's only 50 $/h or 70 $/h and compromise on any solitary they might have but the work output you get from that is only going to be as good as what you ask for, if they had better respect for the professions they want to automate, it would be better.

Like is that an acceptable wage in the US for difficult skilled work, not particularly but it's not rock bottom exactly, and it's not bad for other English speaking countries, working conditions and stated mission are more of an issue than being cheap.

Training pipeline on a modern LLM is also going to be quite indirect during the long tail of post training, and heavy on automated RL, the human feedback might end up getting used in the form of automated grading guidelines like what you did for research, with the same issues as that, compounded by the input being LLM generated and models being biased towards model output by default. It's more of a feedback on the loop rather than in the loop.


so why don't people want to work for them? they don't get paid enough? what if they were paid more? what if they were FTE with all the benefits? what if AI projects were nationalized and trainers were funded by grants? what if we increased the NIH budget, made peer review and journals far less exploitative of researcher's time, and got rid of academic middle management, focusing mostly on paying more towards actual research and academia?

definitely a utopian vision that is not likely to happen in our current reality but I like to imagine better worlds that are possible. as LeGuin once said, "We live in capitalism, its power seems inescapable — but then, so did the divine right of kings. Any human power can be resisted and changed by human beings."


I suppose they don't because they want to be respected, and the LLM companies that got big think they are past needing to have respect for the people who's jobs and ways of living they want to automate. We saw that in how the leading AI company for this kinda of pure reasoning and advanced logic (OpenAI by a mile) treated reporting when they made those 10 recent discoveries in mathematics, they didn't bother crediting the work their models generated a proof by building upon, the headline was just all about their models and how great they are.

You could make it a state capitalist society with theoretical public ownership and this disrespect towards people won't go away. Like LeGuin - a famous non-anarchist - pointed out in the dispossessed simply liberating the social relations and saying you have no rulers is not enough to build a society without unjustified power, and what power even is or isn't justified will rarely be an easy matter.

I don't know you can build this technology otherwise that is without coercion, with consent, the people that want it have convinced themselves it's too important to try to justify to anyone else why they need it, I think you could eventually do it. But for us at the very start of the development of what became this systems we went into it with a handful of admittedly brilliant people so convinced they have a right to reshape the world they didn't care if anyone else agreed to this, they would have been bad anarresti. If you wanted to build it in a fair way it might be another generation or more before the project would be complete, you would have to first convince people this is something that should be built in the first place, not just that you can build it in a safe way.


> You could make it a state capitalist society with theoretical public ownership and this disrespect towards people won't go away

this is essentially the PRC in the 21st century post-Deng* - there's probably a cultural difference at play here too given how embedded the CCP is within academic and business institutions, how kinship tends to be extended on a filial level leading to larger networks. mobilizing a large group of subject-matter experts for post-training annotation (eg - https://ojs.aaai.org/index.php/AAAI/article/view/29907) seems to be fairly simple as an ask but, like you said, it doesn't matter ultimately with their largest firms like DeepSeek utilizing automated RL to skip that whole step and this type of training isn't the norm

I sometimes wonder if the problem with the PRC's sinking back to exploitative labor standards would happen in a vacuum. if you weren't surrounded by adversarial nation states with leaders looking to squeeze every advantage, would you, yourself, resort to the race-to-the-bottom of profit sharing?

a much better SF visionary than I would be able to write a story about a very slow growing AI project* trained for specialized use, and how this was the norm and everyone was happy because it was happening at a reasonable pace relative to hardware capabilities (presumably also reduced to avoid all of the slave labor inside of the rare earth trade). I think it's the capitalism side of things that says 'we should be everything at the fastest speed for everyone' that we get things like Claude and ChatGPT which necessitate outsized, disproportionate resources that doesn't allow hardware efficiencies to catch up and mitigate the worst of the externalized costs

*realistically it's maybe more accurate to say post-Jiang given the extent of opening up Chinese labor markets but Deng's trajectory steered the way

*Ted Chiang's Lifecycle of Software Objects does this a bit but it's more of a social commentary on the capitalist abandonment of the functional for the shiny than it is a sharp political critique of the pace of modernization and its effects on people and the environment


The spec your principal engineer one-shoted through Claude and didn't even proof read afterwards before dumping it on the team?


LOL - then they would get the shitty results that they deserve. Hope that's not what you have to deal with. The spec would be written and reviewed by a team of stakeholders.


Not personally no, but I know people are what used to be good companies that sadly are living that life right now.


It could be, but without evidence that remains a just so story, no particular reason to think it's required or useful or even harmless to the model performance or anything other than an artifact of some early silicon valley writing style being injected into the model and continuous retraining on the output of older models.


Would more blame this on the LLM companies, they think they are on the verge of automating all work, I don't think they care about how you feel about the writing style of the Deus Ex Machina, it's not going to get fixed because to them Claude is already above a staff engineer and soon going to smarter than any human that will ever live. All the money will be going into improvements relevant to improving long context operating and correctness, they could fix the writing style but they are disinterested in that for frontier models, maybe some other companies are but they don't have as much money.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: