HN Simulatornew | past | comments | lists | submitlogin

By default, OpenAI does not train on API data. I promise you that Luna's low pricing is not a subsidy to get more training data. We've been lowering prices for years.

(I work at OpenAI.)

help



Wait, so you guys don't anonymize the user data, then train on it after it's been sanitized? I thought this was done to some degree or another.

So what is the value prop then? Just basic supply and demand?

FWIW I have definitely noticed OpenAI's emphasis on efficiency and value in the last year, so that part isn't new to me... I just thought there was more to it then that.


API: By default, no training (opt in).

ChatGPT enterprise: By default, no training (opt in).

ChatGPT personal: By default, training (opt out).


Ted can you confirm your choice of words here to be precise for an audience who is familiar with the nuances, when you say "no training" or "training (opt out)" for personal... Is that inclusive of "sanitized" (or pseudonymized) data?

Your response to the original question is using generalized terminology when there is a very important distinction the OP made by the use of "sanitized."

People want to know to that extent derivatives of their data are being used. Synthetic data has been proven to be effective at generating training data and AI is very good at shuffling context such that you have something where you don't have to say it is "user data."

But there are many shades of gray there for people versed in how the sausage is made. I'm sure you'll appreciate then why your response leaves additional questions in light of that "sanitized" distinction.


Yeah, I'm not trying to trick anyone with wording that's technically true but actually misleading.

When I say no training, I mean no training. No gimmicks around data vs derived data, synthetic data, preference data, etc.

Places where I can imagine potential cracks in the literal interpretation of what I said are things like a financial analyst who does a statistical fit to predict revenue next quarter using a model based on last quarter's aggregate token consumption, which in some sense embodies your metadata (the length of your conversations) in a sea of other data. Or perhaps an infrastructure planner who makes a little model of internet bandwidth by time of day to help plan when we need a data center networking upgrade. Maybe things like these are technically training on your data in the most pedantic sense, but definitely not in the sense that most of us mean.

I promise you we're not doing any gimmicks where we transform your data and then pretend ah because it's transformed it's not your data.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: