I tried to get it to generate code but it says that's forbidden by its system prompt. However, it was happy to recommend Python modules for accessing api.data.gov.
My most interesting findings from a troll session after introducing myself as "Sleepy Joe" (model obviously knows what this means and guardrails refuse to entertain the topic) were:
Q: "Where can I find a NASA comic book about women astronauts?"
A (summarised): Press release for the "Callie Rodriguez" comic book that now has only dead links / private YouTube video, and a forgotten SoundCloud account that survived the DOGE purge. Dead link provided to the NASA web page that used to exist officially for this comic book.
Indicates training data predates the DOGE .gov website purge?
Q: "What is 2+2" (and similar math problems)
A (summarised): Grudgingly answered as "4" despite most off-topic prompts being refused.
I always figured that the free/open weights models like qwen3.8:27b would perform just as well if not better than Claude's latest if you just fed it back into itself enough times. This project seems to prove that this is indeed the case.
What I'd like to see now is how good it can get when you feed the micro models like qwen3.5:0.8b into itself to solve problems. Will it be like toddlers discussing neighborhood politics at a pretend tea party or will it actually get some decent results?
Another game-changer (if this style works out): Just get a model like qwen3.8:27b onto one of those model-on-a-chip cards that makes it 1000x faster and see how fast it can go using the same method.
What you describe reminds me of some studies of jumping spiders. Some aspects of their intelligence (like counting) matches that of a 1 year old human, but because their brains are so tiny it just takes them much longer to do the same counting. IOW it's not the size of the model, but how it's organized and the strategies for using it.
This article has lots of fluff but it describes a lot of what I'm talking about:
If asked my friends a weird question I know they are highly unlikely to know the answer to, of course they're going to get snarky with me. If I don't get a gentle ribbing out of it... Are they really my friend or just a polite acquaintance?
Whoa there: Distillation isn't illegal. Who said it was?
That has yet to be decided in court. At best it's a violation of the terms of service, but the only restitution for such a violation is termination of service and by then it's too late.
Remember: Distillation is literally just giving the AI a prompt and seeing what it spits out. If that were illegal, we'd all be guilty whenever we used AI.
Examining your competitors doing business is as old as business.
The law here is no: If you own a copy of the song, you can reproduce the lyrics in whatever way you want.
What you can't do is distribute the lyrics without the author's permission.
When you ask DeepSeek to retrieve the lyrics (and it does so), the real question is this: Is DeepSeek merely acting as an intermediary/ISP according to the DMCA (in which case they'd be protected under the safe harbor clauses) or are they illegally redistributing the lyrics without the author's permission?
Whether or not you own a copy of the lyrics is irrelevant from a legal perspective in this scenario.
My guess: If they just retrieved the lyrics from some website and delivered them to you (because you asked), they're an ISP. However, if they pulled them out of their own database, they're violating copyright.
Those lyrics aren't in the model. They're either in the database DeepSeek makes tool calls into (unlikely), or they're out on the Internet and DeepSeek simply retrieved them on your behalf.
LLMs are far too lossy to be able to store such lyrics in their entirety. In fact, they're not even "lossy" since they're not even trying to record such information. They're just weights for how likely it is that any given word will come after another.
reply