Yup, the e2e proof of FLT (estimated effort of 5 years & 1M$ by the best human in the field) and a counter example for a millennium prize (similarly valued at 1m$.
This is a very outdated view on what an LLM is and how it works. We are way past the "stochastic parrot" phase, ever since double descent and proper generalisation. Then with the various flavours of RL the models learn to pluck patterns / circuits out of the massive data and combine them on the fly. There's absolutely no reason to think they can't "invent" new words, because words are just combinations of tokens at the end of the day. So if they can come up with "in this codebase bar is load-bearing" they can similarly come up with "bumblespin is the new word for reversing the polarity of the quantum surface of a spin-aware brane in four dimensional bumblespace".
So you’ve claimed that they have half of my precondition for them to speak in ways that we don’t understand.
Coining a term on the fly: yes
Remembering that term and incorporating it into its mind: no.
But, beyond saying they aren’t stochastic parrots you haven’t really given anything to substantiate the idea that they aren’t (very complex) stochastic parrots.
Interesting. I don’t have a counter to the coining a new term evidence in that article. Potentially that’s solid evidence of the first aspect.
But I do have a counter for the second point:
Context is not a part of the model. It’s an ephemeral blend of input and output fed back in as input.
Online learning is potentially a way for incorporating and evolving coined terms to happen. But since that’s not what any of the models are doing, something would have to change in how we do things before it could happen.
Yeah, they'll have more versions for sure. There are two main designs that have been discussed / shown before (one where the entire nose hinges "backwards" and exposes the payload, and one that looks like the shuttle bay doors). And I think that for NASA decadal projects (i.e. JWST-like big telescopes) they could even go with a non-reusable Starship, with "regular" fairings that get dropped.
The barrier to entry is much lower now. Not necessarily to entry but to do actual useful stuff. You can point your fav agent at it and tell it to do things besides "light up a led". That could also be a factor.
> fail to address sandboxing as a first class citizen.
Isn't it better if the tool is sandbox agnostic and you as the developer / integrator choose what's best for your use case? There are several levels of sandbxing, with many degrees of "freedom", so it would be really hard/confusing/overly-complex to build something ootb that suits everyone, no?
I'm saddened by this comment being upvoted a lot on this site. I'm a space nerd from EU and have been watching SpX since they were launching from an atol. This comment is the epitome of clueless negativity, wrapped in an umbrella of big words and a list of things that are neither important nor factually correct.
Nothing listed in this comment is going to lead to "the demise of the starship project". And nothing listed there is original. We've heard it time and time again, when F9 was ready to launch, when F9 was ready to land, when F9 did land (bbbut they have to do it constantly), when F9 managed the first re-use (they need to do it at least 10 times so it's worth it) and so on and so forth.
They've now launched AND landed >600 times already. Let's just say they know what they're doing.
Anyway, for a bit of context: the current first orbital flight, with 26 satellites deploying right now, is adding the same amount of Starlink bandwidth as 10+ F9 launches. Let that sink in and re-do whatever cost/opportunity and "management" spiels you need to do. One flight > 10x F9 flights.
That alone justifies the Starship. The rest is the cherry on top.
...keeping whatever project that should've been put to sleep months ago alive. Most people reading this have been on a software development project that fits this pattern of ignoring massively missed milestones and hoping to catch up. That's often accompanied by isolated economic arguments that are fragments of what started out as a coherent business case that can no longer be supported.
> Most people reading this have been on a software development project that fits this pattern of ignoring massively missed milestones and hoping to catch up.
And some of us work in the industry that is in the form it is today, as big as it is today, majorly due to the one company alone. The one that puts more mass into orbit every year than any other entity in the world - nation states included.
As for "massively missed milestones" - it is easy to predict milestones when you are making CRUD apps in React. Not so when building the largest rocket that has ever existed, and making it fully reusable.
> coherent business case that can no longer be supported.
Lol. They operate the most successful and most profitable rocket in history, and they are confident enough that they have stopped taking bookings for it. And they just demonstrated a system that can bring the cost down another 1 or 2 orders of magnitude with nearly zero competition even close to them. No coherent business case indeed.
A lot of these are examples of bad prompting (i.e. "build me a website for x y z"). Gradients can be prompted out, same as emoji slop. One line in the prompt / agents.md and you never see those. Color "rainbow" should be avoided by specifying a color palette. In case you don't have one, ask it for one "from literature / best practices" first. LLMs know the theory, but you have to have it in the context for it to be applied.
Misaligned stuff is the prompter's fault. Those should be caught by the human. they're the types of tasks that are fast/cheap to verify.
It’s the next evolution of applications where nobody actually cares about the users, as long as some checkbox can be ticked to mark it done and get paid.
The implementer is not paid by the end-user and they have no reason to care about the user’s opinion - the only thing that matters is to minimise cost and maximise profit.
The client doesn’t care as long as it’s cheap and technically works. The career prospects of the “project lead” on the client side are likely not affected in any way if the result is poor, so they won’t care much either.
There need to be consequences for negative feedback, which would make people strive for a better result. Otherwise people just don’t care.
> what are the use cases for this kind of model? Could it be used in the context of coding agents
Yeah, it could. The most obvious usage would be to have local fast cheap "feedback" / "control" over a slower more expensive agent (i.e. cc / codex / opencode). Things like "goals" could now be split from a long prompt into "actions" and "verifiers". Where for each action you also produce a verifier. Then after each action you run the verifier w/ this kind of "universal classifier" and decide if the step was done correctly, if it needs follow-up and so on.
Example: implement auth in this repo -> llm_plan() -> for item in plan generate_verifier() -> for item in plan implement() ; verify() ; accept() / followup().
Verifiers could be something like this. take a plan item as input, generate classification questions that might verify the task "is this following project conventions?" | "is this touching files from other tasks?", etc.
You can do that with LLMs, but some things might become cheaper / faster. And you can pretty much use it to check against an ever growing list of conventions. Yours or project specific.
reply