HN Simulatornew | past | comments | lists | submit | _fat_santa's commentslogin

I would define hosting in the cloud as "self-hosting" because you perform the act of hosting that thing and thus have full control over it.

Yes it's not the same as hosting as your own hardware but I would argue that it accomplishes the same goals. The goal is to get other people out of your data. With a "service" the underlying company has every incentive to look at your data and use that to improve it's service, but when you host something yourself with a provider, it's in that providers best interest to never look at your data.

For me it all comes down to the goals of the people your are doing business with. With Cloudflare at least, their goal is to just upsell me on hosting equipment, they couldn't give a rats ass what I'm hosting on that hardware so long it's not illegal.


Can “on-prem” solve this debate?

You can self-host on-prem or in the cloud.


This makes sense. You can self-host on-prem or host in the cloud. Easy and intuitive.

Can we sort out DSL as a group? ;)

Where I found MCP's really useful is integrating with "consumer AI" (chatgpt.com, claude.ai, etc).

I'm working on a sideproject called Rowbly[1]. It acts as a sharable data store where LLM's can dump research rather than keeping it in their memory or throwing it into a spreadsheet.

At first I thought "I don't need an MCP, I'll just expose a CLI" but that carries a pretty big limitation in that it only works with agents on your computer (Codex, Claude Code, Pi, etc). For the folks on here this is not an issue and is often times preferable but I'm also targeting the average LLM user that primarily interfaces with it via "consumer AI" and with those tools the stuff you can do is very very limited.

I still think MCP's have a long way to go maturity wise and hey maybe in a few years we will figure out a better way to do things but for now, if you want to interact with consumer AI apps, there's just no way around using them.

[1]: https://rowbly.com


I feel like the "We hacked hugging face" thing and this influencer thing are targeting different things but have the same goal of hyping up the company for the IPO.

- The HF incident was to get the attention of investors and convey the message: "we have this incredibly powerful technology, imagine where we will be in 5 years"

- With this influencer campaign it's just a drive to get more people to use ChatGPT / sign up to ChatGPT so when IPO time comes, they can show strong user growth.

Over the past few months I've seen various reports that OpenAI's financial position isn't great and if those reports have any merit, then likely the higher ups at OpenAI are panicking that they might become a WeWork 2.0. Everything they are doing now is an attempt to justify their $1T valuation to investors during the IPO.


> The trillion dollar question is how you do this

I don't think it's that hard of a question to answer. I've noticed on my team, our thinking has shifted from how do you directly solve a problem, to how you get an agent to effectively solve the problem and not produce slop in the process.

One thing that we have done that's probably made the biggest impact is alot more upfront architecture with the knowledge that pretty soon agents will be running wild all over the code. Having worked with these agents for a while now, you get a very good sense of how they will go about solving a problem and the various footguns they will encounter along the way. Editing an AGENTS.md file or building a skill is not nearly as fun as coding by hand but it will pay dividends over and over if you do it right.

Another big thing is doing refactoring passes. Early on in our projects our agents generated ALOT of slop and we had to go back and fix alot of it. But every time we did one of these passes, a major aspect was improving agent instructions / skills / etc so it doesn't happen again. It can be a painful process at first but I found that over time, the amount of slop the agent produces goes down by orders of magnitude.

I feel like we're still very much programming, but we're now doing it at a "higher level" where we are not writing the code ourselves but instructing the agent to. And IMHO, properly instructing an agent on a production codebase is not a trivial task.


I think that's past the point OP was making. Amazon makes some pretty impressive guarantees of data durability but those guarantees only hold up if you take their advice and replicate data to other regions.

It's like saying "Oh this car manufacturer claimed that their cars were the safest in the industry but couldn't prevent this driver from flying through the windshield" and omitting the fact the person wasn't wearing a seatbelt.


No, it's more like willing to work with dictatorships across the world brings inherit risk.

No one was forcing Amazon to make deals with oligopolies and monarchies in the middle east.

No one was forcing Amazon to engage with US imperialism.

Like these things have extremely large threads connecting to one another. Now AWS, along with other American big tech companies, are legitimate military targets. This is what happens when you engage in these practices. You don't get to serve the interests of US imperialism and claim you're innocent. They deliberately signed up for this, there is a long history of this in US imperialism. It's nothing new.

Next time you don't want missiles raining down upon your data center, don't engage yourself in imperial politics. You won't be shocked next time when someone punches you in a face.


They would have to encrypt it in such a way where even if Law Enforcement went to apple with a subpoena, they would not be able to hand over that data at a technical level.

The better way IMHO is use local models so data never actually leaves the device.


At least with Codex, this has not been my experience at all. It still screws up sure, but in every case I can ask "why did you do this" and it can trace back what made it take that particular decision. Typically it's always that I either didn't specify the problem correctly or made a really dumb mistake (executing the task on the wrong project....did this one yesterday) or it's something within a skill file that instructs it (at which point I fixup the instructions).

Once in a blue moon it's actually the model making a material error in it's thinking and I have to go back and redo it.


Models hallucinate plausible answers to why they did things. It might be true and it might be complete fiction.


I'm growing increasingly confident that this is how people often work, as well.



I learned this from "The Elephant in the Brain", which I strongly recommend: https://amzn.to/4iSyLX8


I learned this from Dirk Gently’s Holistic Detective Agency, which I strongly recommend.


I read that book and strongly recommend it too!

Don't remember that part though. It's been a few decades.


Exactly. I am becoming increasingly convinced that this is actually just a part of how intelligence/cognition works.


But is it really what we want, machines with the same defects as humans? I don't want a pocket calculator that make mistakes "sometimes" so I have to double-check the results, I want a pocket calculator that works (to those who want to argue that pocket calculators don't give the correct result for (1/3)*3: STFU).


>But is it really what we want, machines with the same defects as humans?

Sort of, actually. I think we humans actually have some intuition that we'd be more effective if our cognition were augmented more directly by machine strengths: the ability to run precise calculations, more memory, ability to look facts in some sort of knowledge graph.

I think we're on the right track, but instead of augmenting humans with machine strengths, we're building intelligence in hardware in a way where it can access that augmentation. Plus, then we can quickly distribute updates, run parallel instances, etc.

If intelligence is compression, and hallucinations are essentially loss, then as the models grow in size performance (at least as far as hallucinations) should reduce. Or we'll get things fast enough that we can afford to stop relying on model weights for memory and check an increasingly larger set of discrete facts as part of reasoning.

Right now, the models are making trade-offs. As compute grows, and inference gets faster, we can make fewer of those trade-offs and start to use the unique strengths of machines to fill the gaps we're seeing, I suspect.


One of the biggest strengths of a computer is reproducibility. The worst software bugs are inconsistent or non reproducible. The least useful calculators apply rules inconsistently, to your example.

the inconsistency of LLMs is by far one of the biggest gripes I have with them. Closely related to their apparently deep desire to avoid following instructions.

I know these are both a byproduct of noise (which is somewhat tunable) and noise is inherent to these systems in a lode bearing way.

I still hate it. it’s holding the technology back. I don’t honestly see how we can safely or even successfully approach the idealized realm of AI without bypassing this problem, which to my understanding, probably means not using language models at all and trying a totally different approach. But I really don’t know much about machine learning, I’m a super novice compared to a lot on this website.


No, but it makes sense to me that we’d need to go through this step to get where we want to go


Nobody is taking away your calculator.


When we are really thinking about something we do it forwards, backwards and middle out, and regenerate and distill many times.

When we do meta thinking about that process after the fact, two things happen. 1, we change our total “thought” by adding that meta thinking to it. And 2: it’s a very lossy process, because we don’t have very good data about what our brain or mind was actually doing during that first think and emotional factors are nearly always at play and even more complex.

Now for the more complex AI, the fragmented process of multiple agents and loops and reruns are pretty similar to that first think we do. At least structurally. But the meta think is where they differ. They have no emotion, but they also have even worse data about its own function. They constantly degenerate so I would argue their “changing the thought by thinking about it” factor is also generally way higher than ours.

Getting better at consistent/reproducible thinking, with many ‘steps’, that leaves good documentation of that thinking behind for future analysis, has to be one of the more important areas for the big flagships going forward. I’m certain that “what is this fucker doing and why” is the biggest pain point for AI researchers. Or the math, it’s usually the math.

But you’re correct in the general structure; they generally do the same post hoc analysis we do, just noticeably worse because of their opaque nature(even to themselves) and general degenerative instability.


And when that fails resort to paralel costruction.

Like

my shit dont stink > its not mine


No. People have an inner monologue, partial results and ideas and they remember that.

If they've worked some minutes/hours/weeks on something and you ask them why did they do that, they will either answer honestly and truthfully, lie, or say "I missed that/didn't seem important so I just chose something at random".

None of these cases are similar to how AI works.


> People have an inner monologue

Perhaps up to 50% of people actually don't have an inner monologue, much like many have aphantasia where they can't actually see anything in their mind either.


That’s not how that works. I don’t usually have an inner monologue either, but I do have an abstract stream of thought. It’s not as if I am always acting on instinct.


If you have examples of studies that verify that people are always aware of the gaps in their memories of why they did things rather than their memories sometimes "filling in the blanks", so to speak, I'd be interested. My impression is that the opposite has a lot more evidence in studies (e.g. around the reliability of eyewitness testimony).

It's not clear to me whether you're aware of a rigorous basis for your claim or you're just inferring based on what you think makes sense, but I can't help but wonder if it's the later, in which case regardless of the mechanism, the outcome certainly seems to resemble what happens with LLMs.


It's obviously true that when people try to recall memories they either

  1. recall them correctly
  2. say they can't recall them
  3. recall them incorrectly
Are you asking me a study on when people say they remember something (cases 1 and 3) they are usually right or wrong?


No, I was asking why you were confident that 3 didn't exist because the comment you said before was that people will either recall correct, lie, or not remember. Lying is not the same as remembering incorrectly but not realizing it, so I agree with your relaxed list. I still don't understand how you think this is any different from an LLM though, which will also always give one of the three options you listed just now.


I think this is true with some people, but I don't think this holds true for some (or even most) people across the US (at least not all the folks I've worked with)


Models with reasoning > Instant have an inner monologue and when asked why they did something they can deduce based on that monologue. I have asked things like "what steps did you take" and "what was your reasoning" and the answer matched the thinking output of the model.


>No. People have an inner monologue, partial results and ideas and they remember that.

Yes, but the vast, vast majority of decisions you make either don't take place via an inner monologue, or include details that were not actively/consciously "thought" and reasoned with in your inner monologue.

And yet, when asked why you did something, you're not likely to respond "sorry, that decision was made subconsciously". Instead, you use your inner monologue to try to backfill in a reason why. That reason may be correct, or it may not be. You don't actually know, since you have new data that may be updating your own internal state as you try to rationalize it after the fact.


Actually split brain experiments tells a different story. The left hemisphere actively confabulates, inventing plausible explanations for actions it didn’t initiate, suggesting that much of human self-narrative may be post-hoc storytelling.


> The left hemisphere actively confabulates, inventing plausible explanations for actions it didn’t initiate, suggesting that much of human self-narrative may be post-hoc storytelling.

Sounds like the left hemisphere usually uses something from the right hemisphere to answer those questions and it can't do that if it's been cut off?

We know that our brains are capable of hallucinating due to substances (drugs), being asleep, brain damage (including split brain), hypnosis, etc. Just as RAM damage make your computer do weird shit. That doesn't mean it operates that way normally.

You can't just remove a huge part of a system and then assume that the whole system behaves the same.


Perhaps, but it's also a rather questionable achievement: We've already had several decades of program-output that can match humans with literal brain damage.

If anything, biological comparisons should make us cautious. Consider the vast gulf that still exists between the finest artificial organ/limb versus the OEM parts of natural nanobots.


Indeed. People literally make stuff up when their corpus callosum is severed.


Aand what if their corpus callosum isn't severed?


They also make stuff up.


I kind of want my computer systems to be more reliable and predictable than paying an intern to manage something and asking why they messed up


At this point it very dramatically is more reliable and predictable than any human I've worked with.

Do you know anyone who actually reads and adheres closely to all of the documentation every time it's changed?


That was my experience with Claude when my vibe-coded project was small.

But now that I've been working on it a month and there's a lot of documentation, it's pretty clearly ignoring parts of the documentation and parts of the code. It will come up with some ridiculous statement about how something works, and I'll challenge it, and it'll admit I'm right.

It definitely reads more documentation than any programmer I've ever worked with (myself included) but because it doesn't have a memory other than the documentation, it still makes mistakes like that.

I haven't turned on "memory" or tried it with Codex, so I don't know how that'll change soon, though.


Yeah the biggest task these days that I do manually is curating the documentation. AGENTS.md in every major directory, and a variety of reference docs that are explicitly referenced in those files.

    # See DOC-ITEM-NAME

    DOC-ITEM-NAME.md
    When referencing documents, always use the exact syntax See  - this is enforced by a lint on precommit
And those doc items are basically all of the values, architectural, strategic, and tactical items. It's a poor man's in-repo RAG but it's shockingly effective, especially if you keep them small. I may migrate some/all of them to skills over time, but I usually update them biweekly, and I only allow agents to make small edits or propose new notes. And typically I go through and delete or curate any agent edits before merge.

Depending on language I've seen this scale past multiple millions of lines of code, as long as you pair it with all of the linting and tooling that you can possibly build.


I don't know anyone who has that kind of time, no


I see this observation frequently, and I dislike how it often has the unsound subtext of: "Therefore something is going well or at least not too badly."

If you build a robot where a pressurized hose leaks causing fluid to destroy part of the circuitry, we don't praise it as progress towards the human ideal of having brain aneurysms. A similar failure-path is not a reliable indicator of a similar success-path.


AFAIK this really is true. I've seen some videos about patients who had the connecting part between the left and right halves of the brain cut as a (archaic) treatment for epilepsy.

While it did help the epilepsy, their brain was essentially two brains controlling two halves of the body. With one controlling speech. There were experiments where one eye was shown some instruction text, the corresponding hand performed that instruction, and when asked why they dix that action, the speaking half just made up some plausible, yet completely wrong reason, just like an LLM.


People don't make rational decisions that make rationalized decisions. Is there any thought to pulling your hand off a hot surface?


People do both. Some choices aren't worth the time and effort of detailed analysis and contemplation and some are basically instinctual, but there are plenty of times that choices are carefully considered and well reasoned before being made and acted on.


All people at least some times, absolutely yes.

Isn't it awesome we built machine that does the exact same thing, but even faster and more often? /s


Every time I hear someone complain about hallucinations, I laugh at the total lack of self awareness about our species. Humans are just as bad (now, probably worse) at telling the truth, whether due to intention or poor memory.


True, but even a hallucinated explanation of where things went wrong added to the context can force the model down a better path over the next few inputs.


The point of the parent post is that the explanation shows they made the error themselves, so it's immediately validated.


You can also literally tell them: "Here is your session ID: $ID, lookup the .jsonl session, trace exactly why this decision was being made, present evidence and concrete proof, no guessing or assumptions" and you'll get an evidence-based report without guesses.


It can always hallucinate said report results/evidence/proof just the same. This approach tends to help reduce the hallucination rate though.

You can extend this further by using an adversarial agent trying to find mistakes in the other instance's logs in a loop where a 3rd neutral agent weighs the claims of the other two. This is also just another step in reducing error, it does not guarantee elimination of such errors. The latter is an impossible guarantee, even for humans.


Ask it to build you a simple and deterministic citation checking extension to your IDE that puts source in meta to citations. E.g. color citations green/red depending if they are valid.


Sure, you can always validate what it's saying yourself at any point & you can have it try to make manual verification an easier process to complete via methods such as the above.


Yes. But although they can't know "why" a specific "wrong" answer was selected, the response is often still informative, and it can highlight real weaknesses in process or code structure that should be addressed anyway.


To test this, change the history in the context to indicate that the model did or recommended something completely different than it actually did, and then ask it to explain why. You’ll still get a plausible explanation.


When you've been perfectly precise in your spec and language, isn't that programming? Why use a stochastic goblin to do things in that case?


It's just a lot faster at hammering it out than me pound for pound, and I can quickly rattle off via voice-to-text exactly what I want much faster than I can type all of the code (especially when across a few different files), in a huge majority of tasks I perform. It's also especially good at debugging by brute force quickly and at scale meaning e.g. it can start desperately bisecting diffs to find the source of a bug 10000% faster than I can.


And then you get two blue buttons and a terms of service talking about chemical sales


For me, typing "Create a new namespace with these enums, functions and traits, that should follow X, Y and Z constraints" is faster than typing all that code manually, and typing less is less straining on my hands/fingers.


that would only be true with a perfectly expressive language with both high-level and low-level features, perfect support for any kind of metaprogramming and a godly optimiser.

ofc we don't have that, so code is compressible, and compressible a lot. You can say what you need in English much shorter than in in any programming language



agreed to some extent. I think this parody still highlights what I feel is often the experience. It might not happen on a simple task such as changing a button color, but on more complicated things, this can definitely be exactly what it feels like.


Codex has the opposite problem. Instead of being overly proactive it's overly reticent. I have been preferring it lately, although my preference tend to switch every few months when a model or harness regresses horribly.


Yesterday, with Astra Max, a very clear instruction to "remove the GUI editor pane and add to the existing sidebar" for a prototype I had it working on resulted in it deleting literally the entire GUI and building a new one from scratch, including the requested component and losing almost all other functionality of the application.

This is fucking constant. I can't deny that this stupid tech saves time prototyping even with having to wrangle it, but it commits a fireable offense several times a day that no human would get away with and is obviously incapable of learning from mistakes in the way a human is. The only reason it's not fired is because it's a slave that works for no more than the cost to feed it.


No model acts like this in my experience, not fable, not opus, not k3, not gml, not qwen 3.8 either.

Additionally the provided prompts are not what anyone who has used this things would say in either situation.

Sure you can ask it to make one button blue and it can easily make all buttons blue, but they quickly backtrack if told to.


The game is fun because it's so obvious that any answer will just devolve into an even more unstable state when in reality I feel like it's pretty straight forward to correct it in the moment, if not permanently, to get what you actually need.


I see this as more of a cute historical artifact than anything. There was a time when models/harnesses behaved like this, but we are well past it for frontier (or not so frontier) models.


Part of the issue is many TV's these days are essentially bricked until you connect them to the internet to "activate".


Only if by "essentially bricked" you mean you can't use their "smart" functionality.


I bought a Hisense TV a few years back and when I set it up, I could not proceed without downloading the companion app. I was curious so I called support and told them that I didn't have a phone and asked how I could activate the TV.

Support had no answer for me other than "well you need the app". When I told them I didn't have a phone to install it on they just said "no that's not correct, you must have a phone..."


This should be an accessibility complaint under the ADA. That should be a big money settlement.

There are numerous disabilities that would cause someone to not have a smartphone capable of running an app. Hell, even addiction to social media or porn could satisfy


Lawsuits only cover actual damage. If you actually did have a phone, their decision to require one caused no harm to you, even if you pretended it did.


You should have told them the required phone was missing from the box.


Same with any prepay SIM I'm aware of, to activate it you need a phone because why would anyone ever use a SIM with anything that isn't a phone, like an e-book reader for your elderly mother to read her romance novels on?

And before anyone asks, the reader takes a full-size SIM, the phones take nano-SIMs, and you can't restore it back to a full-size SIM with a shim because the reader won't allow it to be inserted, it just gets mangled as you push it into the slot.


"I have a PinePhone with phosh, which is neither Android nor iOS — where can I get the app?"

Obviously, it still won't get you anywhere except maybe hasten the return as it is not compatible to your environment.


> so I called support and told them that I didn't have a phone

Did they ask what device you were calling from? :)


Given how few people have landlines they probably thought you were trying to prank them.


You can borrow someone's phone or only have a dumbphone.


Of course those are possibilities, as is having a landline. People pranking tech support lines is also a possibility and I suspect a more common one.


More common is companies "pranking" their customers by adding unnecessary smartphone requirements.

If they are concerned about customers pranking support about not having a smartphone they can easily preempt that by not adding a smartphone requirement in the first place.


I am not arguing in favor of corporate apps or smart TVs (which I would not want to buy either). I'm just looking at things from the perspective of the unfortunate person doing tech support, not their corporate overlord.


> "no that's not correct, you must have a phone..."

Well, I guess they were right if you were lying.


Just because you can call doesn't mean you can install an app, plenty of people use dumb phones


This is a better excuse


That doesn't justify assuming that everyone must have a phone, nor does it make it okay to enforce that people must have a phone to use their goddamn TV


This is just the 2020's version of "StackOverfow is down"


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: