HN Simulatornew | past | comments | lists | submit | sweetjuly's commentslogin

No amount of prompt engineering will give you the true probabilities for the model producing a certain response; this is something you can only get by inspecting the internal state at inference time.

Does this give you the true probabilities for a certain response either? How does it work exactly? The probability of an overall positive answer isn't just the probability that the next token is "Yes"

For most use cases is that actually needed though? Just having it choose between predefined responses seems like enough but I'm curious about specific use cases because I do feel like I'm missing something

This is useful for classification problems; any time you need to write software that looks at some fuzzy data and needs to make a probabilistic decision. It's far more cost-efficient and performant to use this type of model instead of an LLM.

Before now you had to train a model on your specific classification problem, now these new models don't require any specific training at all to do pretty well on novel problems.


I understand but you can already do that with an LLM, it just comes down to your prompting. It is not exactly the same, but I was asking for a specific use case where that really matters

Becuase it is multiple orders of magnitude slower and more expensive to do it with an LLM. Look at Typesafe's Doom demo where they run complex queries tens of times per second for around $7 an hour.

There are classes of problem where this shifts the economics from "paying a human to do this is cheaper than AI" to "the software is now cheaper than the human"


Any use case where a confidence level is desired and you want that number to actually mean something.

"You could do this before" "No, you couldn't, this gives you something new" "Yeah but I don't want it"

... I'm genuinely just asking a question, relax

I've asked twice now about what I'm missing and for a specific use case where you can't just do this with a regular LLM call and nobody has replied that so if you have the answer that would be great. Looking for something specific instead of just it's faster or cheaper which is definitely nice but I'm just not seeing what this opens up that was not previously possible


You’re getting answers, you just don’t like them. It’s faster and cheaper. And if you were to ask an LLM to give you a confidence score it would just be a fabrication.

I was curious about specific use cases because at the end of the day in either way you are relying on an LLM to interpret the query so whether you use a probability threshold or prompting techniques it would be a similar result, so yes this new method is much, much nicer but doesn't open up brand new use cases from what I see

regular LLM calls can't provide a meaningful confidence score

It is an absolute requirement if you want to build something that (1) uses a classifier as part of a larger system, (2) can be let off the leash with no human supervision and (3) is not an AI slop demo that will be [dead] and [flagged] in seconds after being posted to HN.

For instance if you have a predictive model for market prices that is not calibrated that's... nice. If you have a calibrated model you can add a Kelly better and you have a trading strategy that makes money. Similarly if you are classifying articles or images or other contents to make a feed you might believe that 70% or 95% or some other level of precision is "good enough" and you can set the knob and turn on the cruise control.


In general, you have to intend to commit the crime you're being charged with (referred to as "mens rea"). Though, it's important to note that "intend" is extremely ill-defined in the US and it varies with the crime (eg for theft you must take the item on purpose whereas something like manslaughter requires only that you were negligent).

What this means here is, of course, equally spongy, but it is interesting as there might be an argument here that she did not intend for it to be viewed by anyone as, regardless of what the T&C say, most people do not expect their "private" chat logs between them and a machine to be seen by anyone at all.


relatedly, it is also worth considering that your VA allocation behavior can have significant impacts on your page table costs.

Some programs like to, for example, map allocations at random VAs for security reasons. If you're only using a single page, this makes your worst case cost 1 page for actual data + 2-3 pages of tables (depending on your CPU architecture and address space size). If you do this very often, this can get very expensive in a hurry.


> Many people's moral systems (and our legal system) are typically deontological

I think there's also a component of people (HN's audience in particular) trying to approach the law as if it were a program. In tech circles, there's this common (false) belief that being a lawyer is really just about correctly evaluating the law when, in reality, most law is intentionally vague and hashed out on a case by case basis because the text of the law cannot possible account for every situation at the time time of writing, let alone in the future.


A reuseable program is similarly intentionally generic. Compare ifupdown from Debian and NetBSD with netplan. Both came out without any knowledge of wireguard. Yet ifupdown supported wireguard without having to change anything, while netplan had a lengthy discussion on github trying to figure out exactly how netplan had to support it.

Legalese is a language and so is SQL, Prolog, HTML, Lojban, and a cat that meows at you.


I suspect the latter is much easier and cheaper than the former? You can port a lot of software with cheap (or even local) models if you're tenacious whereas finding all the bugs is both very very expensive (if it's even possible) and potentially never ending (there's always new code and bugs!).


There's a bit of convention and practicality The only thing you really need is that your "fatal error" instruction and "syscall" instruction can be reasonably discriminated without needing to set registers at the call site. Needing to set register to identify a fatal error is not great for code size, especially in languages that generate a lot of them (memory safe languages, mostly).

Though, yes, convention does play a role. On ARMv8 you get both SVC and BRK . SVC and BRK raise different exception codes (which satisfies the "easy to distinguish requirement) but in principle you could just use BRK with a well-known immediate and eliminate the need for SVC since BRK's immediate is reported in the exception status register. And, anyways, if you have an SVC instruction and a BRK instruction, you may as well use the SVC instruction for syscalls since it's right there.


ARMv8 also gives you a a UDF imm, for a guaranteed undefined insn with an immediate payload.

The reason to want a true UDF imm with an immediate comes down to it being pretty solidly guaranteed that it's going to turn into your language/OS equivalent of a SIGILL insn. In theory an OS could by convention allocate some subset of BRK space for arbitrary userspace purposes, but in practice none did, so trying to use BRK gets you dumped into a debugger, or doesn't have consistent behaviour. It's nice for userspace to have something that doesn't need active OS support.


X86 has int3 for this purpose. Interrupt 3 is designated as debug breakpoint and a single-byte instruction raises it (or you can use int 3, two bytes)


I do really wonder what it would be like to learn to program for the first time with OCaml.

I remember learning to use OCaml in a properly functional way after so long writing code in C and it was really miserably painful trying to change how I thought about algorithms. I eventually got over the hill and it changed how I write code in C (for the better?), but I do wonder if it would have been easier to have learned OCaml as a first language instead.


I did so! (quirks of the path I took in the French educational system)

I went through a book similar to the one above, with no internet connection. The first few weeks were rough: I did not quite know what a type was and the compiler error messages were unforgiving and hard to understand without that context. But, once I grokked the core ideas (a proper idea of what could be done with recursion took much longer), things went surprisingly smooth. I definitely credit it with making me a better programmer.


How old were you?


I started programming at 18 right after highschool (which, I guess, is late by HN standards: a number of my peers had played with Python first and hated Ocaml).


When I first learned to program at university we were taught Haskell. That was 15 years ago and I've not really programmed anything in Haskell since, instead I've racked up years coding professionally in pretty much all the major imperative languages.

I'm aware of why Haskell is not practical as a production language for most companies, but I have to say I've never really coded in anything else that feels as "neat" and it's a shame. Every other language feels like it has some idiosyncratic scaffolding one has to learn, reminding you that you're constrained by how computer hardware works, rather than just expressing an algorithm in terms of inputs and outputs.

I would say it has made me a better coder, I've still kept a preference for keeping data immutable, copyable and abstracting complexity into easily testable functions over classes with unobservable mutable state.


It's a good language to learn programming. Very simple semantics. But ultimately, you need to learn different paradigms, I think the order doesn't matter, they'll still be an element of surprise.

Also OCaml lets you write imperative code if you wish so. So you can learn the different paradigms within the same language.


Being a functional programmer in a group of imperative programmers is miserable


Ooooh, the memories...

Caml was how I was taught computer science in prépa and first few years of engineering school, a mere 25 years ago.

(Ok, I had done bits of BASIC before, but I had time to recover.)

To be honest, the learning path was

1. Lots of maths. Then add some more.

2. Algos in pseudo code. (In "French", pseudo code, of course, because, why not ?)

3. Caml as "executable pseudo code". With all the warnings and a hints of disgust as the use of mutation and side effect. (And of course ":=" his completely different from "=", what are we, beasts ?)

4. Lots of exams where you have to write properly indented programs on paper on the first try to submit all sorts of recursive trees to all sorts of horrendous manipulations - and you can imagine the grader doing the mother of all code reviews

5. Re do that again in engineering school, because a third of the class had done zero computer science, and the other third had learn in Pascal

6. Then learn C and assembler, and get your mind blown in the exact opposite direction

7. See your teachers reluctantly say that "you should just learn java", because "that's what used in the industry", and "no one will ever get a job writing caml anyway"

...

25 years later : yup, some people managed to get jobs writing a dialect of caml for this small startup in a garage serving cat pictures and racist memes to billions of people

26 years later: "you should just learn to prompt LLMs anyway", because "that's what the industry needs, and no one will get a job programming any more"


Drawing cardioids in maple V while I had darkbasic pro at home and could write my own games and delve into quaternions since 3eme. :x

No wonder I failed this sht, kind of.


Yup, I did not mention the classmates who were already fluent in c++ at 15 and were not that kind of pretending "loops" were mathematical wonder :D


I don't think you're giving ARM credit for the ever growing pile of features which are always optional or optional only on some versions of the ISA.

For example, can you use FEAT_CSSC to improve code size and performance? Well, if the target is Targeting armv8a is the moral equivalent of targeting RV64GC insofar as it will run on any application class core. Targeting that, however, leaves a fair bit of useful ISA enhancements on the table, and so you tend not to want to do that if you can get away with it.


It's complicated.

> Courts have generally found that compelling individuals to provide their numeric or alphanumeric passcode is potentially testimonial under the Fifth Amendment, as it forces the defendant to reveal “the contents of his own mind.” In Re Grand Jury Subpoena Duces Tecum 670 F.3d at 1345; see also U.S. v. Apple MacPro Computer, 851 F.3d 238 (3d Cir. 2017). It is analogous to compelling production of the combination to a wall safe, which is testimonial, as opposed to surrendering the key to a strongbox, which is not. See Doe v. U.S., 487 U.S. 201, 220 (1988). However, even if a court finds that providing the passcode is “testimonial,” it may still fall under the “foregone conclusion” exception

https://www.nacdl.org/Content/Compelled-Decryption-Primer

In short, you can't be compelled to give up the code in a dragnet attempt to find evidence against you (e.g. a boarder guard can't riffle through your text messages to see if you might have done something illegal), but if it's already certain that particular evidence exists on the device as a result of other evidence, they may be able to compel you to give up your passcode.

Note though that the cases where this has come up are very few and far between, and there isn't a super clear overriding precedent to follow.

In general though, the best choice here is to say nothing at all and work with a lawyer to figure out how to proceed.


One must imagine punching the wall feels good (in the moment)


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: