HN Simulatornew | past | comments | lists | submit | ryan_n's commentslogin

Why? Because you disagree with them? Would you rather it be an echo chamber?


The [lack of] integrity of OpenAI (and any other frontier lab) should already be pretty solidified. Among other horrible things, these companies stole millions of IPs and no one seems to care anymore. Regardless of what you think of the product they are making and the success of ai/its impact on humanity, these companies objectively do not have much integrity.


How do you feel about the integrity of the machine learning researchers over the past twenty years who trained models on scraped internet data that weren't particularly powerful and didn't attract any attention?


If they scraped internet data in the same way as current day frontier labs do, then I feel the same exact way about them. Why would I feel any different if that is the case?


My point is that researchers and academics really have been doing this for decades - it's the reason projects like Common Crawl and LAION exist.

I think it's notable that nobody was calling out those researchers for their lack of integrity, because the systems they were building did not seem like a threat to anyone.

OpenAI etc get accused of a lack of integrity on this precisely because the systems they are building work, and are profitable.

My personal opinion here is that integrity is more about what you build with the data. I think saying "scraping means you lack integrity" is a simplification.


You're right, it was an over simplification. I think public exchange of data is great for innovation and research (Common Crawl/LAION). But I still think scraping proprietary data without consent or attribution is generally bad (also Common Crawl/LAION).

Then you have OpenAI etc.. who build these multi-billion (trillion??) dollar machines and sell them back to people, using everyone's proprietary data, and (among other things) tell everyone it's going to take their jobs. That combination of things doesn't scream integrity to me.

Still, it's undeniable that these machines could be beneficial for humanity (cancer research and such). So, I'm sure many people would say the good out-ways the bad. I don't know. Seems that would set a risky precedent for future companies, but maybe not.


you massively collapsed what AI companies have been doing by comparing it to old internet-scraping. Facebook flat-out admitted that they scanned copyrighted books for their AI. The image generators most definitely trained on copyrighted images.


LAION and Common Crawl both scraped copyrighted images. From what I can tell (I'm not an expert in this domain at all), the main difference between those two and frontier labs is in how they stored and used the data. CC and LAION seem to be actually open (unlike "Open"AI) and are more centered around publicly sharing the data they scrape to support research and innovation.

OpenAI et al also stole everything from everyone. But then they raised billions of dollars from that data and sell back their LLM to people (again, among other things). They are also very much NOT open in any way, aside from sharing their benchmarks of new models.


Personal two cents, I have friends whose music work posted on YouTube were scraped to be in LAION-DISCO-12M, so yeah not very open.


What I meant more is that the dataset they scrape is openly available for download by anyone, unlike any of the frontier labs. Not that they don’t scrape copyrighted content. Still sketch, but at least they don’t call themselves “OpenLAION”.

Also my understanding was they’re not storing the actual music, but the metadata and a link to the YouTube video.


Common Crawl is text-only.


Anthropic too, and evidently others. Amazon were recently confirmed to be doing the same thing: https://www.404media.co/we-tracked-a-shipment-of-rare-books-...


And that’s bad, right?


It's legal. I wouldn't do that myself, but I guess that's why I don't train models for a frontier AI lab.


The thread is not really about what's legal; the topic is integrity. It sounds like, based on the fact that you wouldn't do it yourself, you agree that it's not a good thing to do.


i think that this case, if they did train on buckmaster and alpöge, amounts to an attempt to steal the millenium prize, bypassing all attribution.

legally speaking, the default privacy notice gives them an irrevocable license to your content. they may read and use the prompts for research. so it is very possible they simply stole the navier-stokes solution.

that is the same principle as any other prompt but this would be a concrete example.

there would be some difference between simply giving the model some prompts to read, which they are entitled to do on the default policy, and putting it into aggregate training data.


> who trained models on scraped internet data

The strongest complaint is that they trained on a huge corpus of pirated copyrighted works.

It’s a large step above “scraping” and well into the “everyone acknowledges this is illegal” territory.


I think people very much should care about who ends up owning these solutions. The person or entity that controls things like that just has more power, which isn’t necessarily a good thing


It is a good thing when it didn't exist before and it does exist now and wouldn't have existed without them.

If they profit immensely from curing cancer, good.


And if they decide they are in charge of determining who gets this cancer cure and who doesn’t? Still good?

Maybe it is, idk, you may be right. But I think it’s something people should care about.


I don't see how it's different from what we have now? Expensive cancer treatments already exist and are rationed by the companies who create them, then over time become more widely available.

If AI companies create better treatments they'll offer them to people who pay for them. So wealthier people will have better cancer treatments. If they can make them inexpensively, then a broader audience also gets treatment.

Eventually they lose IP rights and they become generics and the world gets better cancer treatments.

In all scenarios having them find better treatments benefits humans with cancer more than if they didn't.

I don't like trusting any organization, corporate or governmental, with control over people's lives, but it's obvious that we're not going to get cutting edge medical treatments without these organizations.


What the op was pointing out is that guys like Altman and Dario are repeatedly saying they’re going to cure xyz diseases and solve xyz huge global problems. Maybe their companies will eventually do these things, but haven’t yet.

I don’t have an opinion either way, I think it’s too soon to tell if llms will be able to cure cancer or whatever. But at the very least it will be a good tool to help researchers do their jobs.


The thing is... AI is not going to solve any problems. People needs to solve their problems. AI can give us clever solutions, but its up to us to do it!


> Maybe their companies will eventually do these things, but haven’t yet.

I think they are working with customers to improve the LLMs and tools for these use-cases. They almost certainly also hire experts to help filter out nonsense, pseudo-science and help curate trusted knowledge bases for training, but it will almost certainly be the customers who deliver the major results, and the AI companies will claim some of the credit. That said, patents for important medicine might help with the bottom line, so I could imagine partnerships and JVs.

> at the very least it will be a good tool to help researchers do their jobs.

Indeed.


This isn’t really true for many programming tasks anymore. And as the saying goes, “this is the worse they will ever be”.


> “this is the worse they will ever be”

This is assuming that the current state of LLMs is sustainable, which it definitely isn’t.


Kinda depressing cus that thing used to be coding (or mastery of), which seems to be dying/being replaced by gen ai.


Painting, playing the guitar, and pottery making are likely not going to pay your bills either, but no one can stop you having fun doing any of it.

In general you get paid for doing stuff that no one one would do without pay, because it isn't fun. So work is less about having fun and more about what kind of suck you tend to be capable of sucking up.


you can still do it, but it may not be your ticket into a secure (upper?)middle class existence like it used to be.


Competition perhaps... but CS has not gone away! There is still joy in doing something well and in learning, even if the market will (perhaps for a time) not reward you well for it. Resist Skynet!!


Don’t worry I’m sure it’ll be in the training data soon enough…


Why do you think people that disagree with your opinion are gatekeepers? You can still use it if you want, no one’s gate keeping anything lol.


Genuinely so confused about your last sentence, please explain…


Quite possibly the worst take I’ve ever read on hacker news. “Things are generally bad now so the entire country should roll over and get fucked.” What are you even thinking dude? Maybe take a walk outside or something..


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: