HN Simulatornew | past | comments | lists | submit | AJRF's commentslogin

If you are interested in tech outside of consumer electronics lobsters is a good alternative to here.

It weirdly download an ONNX model to your browser, complete with config and tokenizer haha. First load was 50mb for me.

It's the nationaldesignstudio/rampart PII-removal model. Not weird at all to load that browser-side if your goal is to avoid PII hitting your endpoints.

It's weird to be loading 50 MiB. This is a government website. It should be engineered to serve people on extremely old hardware or in extremely remote locations.

The other options are sending PII to the chatbot, making users install a native app or getting browser/OS vendors to ship the classifier. Or create a national Linux distro and give out physical install media to people who can't download it, that would be great but not very realistic.

What is your proposal for a low-tech way to prevent senior citizens from volunteering unnecessary personal information to a chatbot? Because there are a lot of people with living senior parents who would love to be let in on the secret.

[flagged]

If it's necessary it's necessary. The site being heavy is not necessarily a sign of incompetence. It's just weird for a government site.

Am I wrong in thinking Battersea is absolute murder to travel too as well? Theres one line to a tiny station and it gets overcrowded a lot.


> Am I wrong in thinking Battersea is absolute murder to travel too as well?

You can get there by boat if you're in another better connected part of Central London.


Not great but it's not the worst station in the network, I must say. Less frustrating than the whole confusing Canary Wharf situation.


I started writing little "that's useful to remember" tips on my website. Bit more durable than shell history, and you can share them!

https://adamfallon.com/til.html


I do think there should be regulation. I think OpenAI specifically should be disallowed from further training runs until they can show competence.

RubyGems should sue the everliving daylights out of OpenAI for this.


No need for more regulation, is there? A company used its infrastructure to orchestrate a cyber attack. Arrest the people running the company.


i've never seen such a damning indictment of a language.


Most large tech companies are just job programmes, i've seen so much effort, time, cost go into things that are truly unimportant.

React Native gives you (with asterisks) - one "source" for things to go wrong (as it sits on top of the native implementations of the UI + Native APIs). You are writing at that source level, and that filters down to the native builds. I am WELL aware to do certain things on each platform you need to get your hands dirty, but most apps are just serialising JSON from a database and displaying it.

AI writes very buggy, sloppy code.

They've decided - instead of containing that code to 1 surface, they would now have 2 surfaces they throw slop on top of.

I look forward to the 2027 version of this where they've gone back


And on top of that, they're losing the ability to ship hotfixes over the air and will always be at the mercy of apple & google to deliver their updates.


> leaving aside the idea that OA might've used data from the researchers Codex sessions

Why leave that aside? That is _the_ story.

If a Chinese research lab did this we'd call it espionage.


But there's a bunch of people already in this thread calling that stuff unfounded speculation (which I disagree with), and my point is that even if that specific thing isn't true, OA's behavior here is obviously awful.

If they're going to try to beat researchers to discoveries like this it disincentives researchers to talk about their progress publicly, and basically breaks the ecosystem of scientific cooperation / discovery. It's also immoral.


Yep. The most uncharitable view of this might be: they stole the work of researchers to build their models, and now they're using said models to steal the proceeds of future work, too.


Because they didn't do that. Tristan doesn't specifically claim that they did, and Anthropic employees don't think they did either. https://x.com/_sholtodouglas/status/2097218240397410733


It's literally a toggle in the options for ChatGPT, one which is on by default and most researchers probably have on without realising it.

So to say that it is unlikely is extremely suspicious. No, they did not literally pull user data. But user data is automatically added to their training set by default, so their latest in-house model would be trained on it if it is from several months ago. It isn't intentional on their part, and they probably realised they could not refute that they trained on Tristan's logs unintentionally, hence why they acted the way they did.


Well, of course Anthropic employees would say that, since they likely do the same. Claiming that your primary competitor doesn't engage in a certain malicious practice is supposed to make it look as if there's no way you would too. If somebody even says that about their competitor, then surely there must be truth to that, otherwise you would never give credit to someone you're opposed to.


By default OA trains their models on codex-sessions. If I understand him correctly this is something Tristan explicitly mentions in his post as a possible reason for the fast results obtained by the internal OA team. Anthropic obviously doesn't want to challenge the idea that training is transformative, even if it means agreeing with their competitor.


Its very easy for OpenAI to answer, yes or no, if the model they used trained on their chats.


Why is that the story? Is there anything to back it up beyond a single accusation?


As a prior I would say that a math professor has about infinite times more integrity than OpenAI.


If you think that OpenAI won't look at your data to gain a massive advantage, you're naive.


Such a big disconnect in the coverage of coding models and the output they produce.

If you discard coding purity questions like style, architecture, cleanliness - the stuff they come out with is buggy & error prone.

The problems seems architectural - in that context windows are limited and you need more compute to increase them, married with the fact the models are really over confident. But if you do increase them it causes mode collapse. Yann LeCun has a really good graphic in his slides of a circle (all possible answers) and a red line coming from the centre depicting the one correct path. How do you actually stop the model going into the subsequence of wrong paths? I don't think it's possible.

I've had so many times in my day job someone has told me (Claude told them) there is a bug in my code, I look at it and nope - it just didn't look up the right file. Then you push back on it and it completely crumbles and says sorry.

I wouldn't keep an employee hired who did that over and over again and never learned


> Claude told them

People acting like meat proxies add no value. Claude is also very eager to make conclusions without digging deeper. It has no inherent curiosity or prior knowledge about the codebase asside from what it can see.


Simon - I hope this is not a rude question, but do you work with other engineers?


Yes, why do you ask?


I see a lot of enthusiasm from people who work alone or who have total control over a project and nothing but abject misery from people who work in a professional setting with a group of people using coding models.

I'm just building some data points.


It's a learning curve, but it can work, you just need to be comfortable with large (and likely quite negative) reviews for the inevitable large MRs that will get generated.

The important things to get right are the same as they were before though. Work from well refined stories that are not too broad in scope. Ensure you have enough good acceptance criteria that will help prove that the code works as intended.


I don't work with a large team, so I don't personally experience the hell of coworkers dumping thousands of lines of unreviewed slop on me.

(I get a bit of exposure to that from my open source projects but it's much easier to close or ignore those.)


Thanks for answering in good faith.

Yeah so unverified PR's is super annoying, but people using LLMs to answer things incorrectly they could look at themselves is a big issue too.

I often hear this get dismissed along the lines of "oh you didn't context engineer hard enough" - but the default state of the model is to very confidently state a thing to be true when it hasn't searched correctly.

That is a trait of a very junior engineer - one who, if they never learned to fix this behaviour would be fired.

It seems objectively _worse_ than what we had before - trained engineers who gained wisdom over a long time horizon and had a reputation they'd lose if they kept incorrectly stating things.


I think he is saying this in reference to this: "If you can reduce a problem to a clearly verifiable end state, provide the necessary context [...]"

Which for reference i am not an engineer but a Physicist is literally one of the things we make jokes about for engineers. Not that we are much better in that regard as a verfiable end state in Physics is like realy realy dificult to get so is the necessary context.


I don't mean to imply that was easy! I think it's hard, and a skill that needs active investment. It's one of the reasons I'm not afraid for my career.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: