Keysmashing my keyboard resulted in 86% confidence that the text was AI written. I don't think this is a particularly good classifier, I've never seen an LLM output "kad jfkhasljkdhf laksjhdf".
Edit: If that's not realistic enough for you, the text "Hello world! My name is GravitasIsOverrated and I like coding and cooking. This text is 100% genuine, and not AI generated at all." results in 85% confidence that it's AI generated.
More broadly, I don't know why this would work. Qwen/Jev/whatever doesn't magically have the ability to discern AI-authored text from non-AI-authored text, and will increasingly get worse at it as the hallmarks of AI-written text change.
By transaction volume yes. Presumably it’s not all spending though, if you regularly win back some of that money. The same dollar is probably wagered multiple times before it is lost.
Ontario lotteries pay out about 58 cents per dollar in, so that's still $13,000 annually out-of-pocket for a total spend of $31,000. That's crazy money.
Sportsbooks typically pay back more like 90%, brick and mortar casinos 90%-99% depending on the game. I don't know what Ontario mandates for regulated online casinos, but I'd guess they're at the high end of that and focus on volume, return custom and dark patterns rather than extreme house edge.
If it's 98%, the average household loses 0.2% of their income gambling.
Two notes - 1, this applies only to a pair of base models in the Korean market (for now) and 2, the info is still there, but it's on the central screen and not on the instrument cluster behind the wheel.
As a driver I would be so freaking confused staring at the empty space behind the wheel. Maybe it is just my Gen X experience of learning to drive on a stick and paying attention to tachometer and coolant temp.
I drove a Toyota Yaris for about 10 years and those cars had the instrument cluster in the middle of the dash. It didn’t take long to get used to and one of the unexpected benefits was how much better forward visibility was. The Yaris wasn’t a big car but honestly the drivers position felt a lot more open and less cramped with that layout, and realistically the center of the dash is almost always empty space anyway. I personally liked it a lot. Not sure the “no instrument cluster + massive iPad in the center” is going to be quite the same feel or work as well, but everything being in the center feels way less weird than you would think.
> Maybe it is just my Gen X experience of learning to drive on a stick and paying attention to tachometer and coolant temp.
You and your fancy cars... my 1981 Vanagon doesn't have fancy stuff like a tach or a trip meter or a temperature gauge. Who needs a tach when you have ears, and anyway there's a colored dot on the speedometer at the top end of first, two dots at the top end of 2nd and three dots at the top end of 3rd. If you go over, the rev limiter will let you know.
While working on the speedometer, I did drive it a bit with no cluster installed, and it was weird, but I could get used to a bit more visibility, although the extra visibility was obstructed by the steering wheel.
I bet you had a separate indicator for left turn signal and right turn signal, but do you really need that? --- a single indicator if any turn signal is active is plenty.
I've made the switch before. At first yes of course you do stare at the dash in front of the wheel for a sec before remembering it's to the side but I didn't find the learning curve too difficult. After a month or so it's second nature.
I still don't understand how these massive touchscreen displays are allowed in vehicles.
I had to rent a van that had everything "integrated" in its display. Absolutely atrocious software, latency and responsiveness for standard things such as temp and fan speed.
Also withholding this "extra" means you have to take your eyes off the road and look down to the right/left to make sure you're doing the right speed. Sounds a bit over the top, but I wonder if these designers think about the fact they might actually be the ones responsible for people dying in crashes.
Kimi K3 was an interesting model only a month ago, and now we're looking at the same performance for 1/20th of the price. Wild how fast this is advancing.
Real question: is there anybody that is both maintaining alpha-dev capability by keeping abreast of all these daily changes, while also reserving enough time to actually work?
Seems like we've reached the event horizon of whether AI advances are worth paying attention to.
I feel like it's most useful to get in a bit of a groove with one setup and then poke your head up every now and again and try updating. Staying at the bleeding edge of anything can be a bit of a treadmill.
I think the play now is to just try out whatever the best new model is every time you see a headline that fundamentally reorganizes your conception of what's possible.
I enjoy using opencode go to play around with a lot of different models. I wind up using deepseek v4 flash for most everything, stepping up to minimax m3 if that doesn't cut it, finally preferring GLM for complex tasks or important planning I want to go right the first time
I recommend opencode or something akin to it to play with models. Any big model updates or hot new ones will naturally run across your desk that way
I don't think you need to be keeping abreast of them really, you just need to be using the best model you can get enough tokens from, which for many people is Fable 5 @ $200ish, ideally fanning out implementation to cheaper models
I would love to be shilling for Anthropic, but I am not. I am part of a group of about 30 developers, and 80% of them are using Fable 5 and very sold on it, with the remainder being committed to Sol. Both are competent, but among our set (who will try anything), Fable 5 is definitely winning. The fucking refusals for security work are insane though, and I hate them.
I use Sol and Grok 4.5 as my inline debuggers/reviewers, and both do well, and are decent at token save. DeepSeek V4 Flash 0731 found some interesting bugs when I tried it a few days ago, and I'm curious to see if that also joins the code-review line up
In my use, DeepSeek v4 Flash (which replaced the quite excellent MiniMax M3) lags behind GLM 5.2 & Muse Spark 1.2 (let alone Kimi K3). Also, K3 is a much bigger multi-modal model, while Flash is text-only and likely optimised for coding tasks.
Yep, and the v4 flash final is about 2.5x slower than preview making it no longer a fast model, in fact slower than Luna and bigger models in many cases.
Spark is actually the interesting one imo. It's significantly better, also significantly faster. If you are ok with letting Meta soak up your data (which DS does too) it's also the same price.
I've been using this DeepSeek model the whole day today after building with 5.6 Luna extensively over the last week and I would disagree, at least for Rust + OpenGL.
DeepSeek just spend almost 2 hours trying to figure out why terrain textures were not working. It tried everything over and over again, it even had reference code for meshes on how to setup the rendering with materials, and it could just not do it.
I finally gave up and gave it to GPT-5.6 Luna instead, and figure out in a single prompt after 20 seconds, that the terrain mesh was being initialized with None in the material slot.
Other tasks it has managed to figure out at least, but it is significantly slower than GPT-5.6 Luna and it requires a lot more iterations.
That's roughly my experience. Luna is extremely efficient and at higher levels of reasoning and longer running tasks more capable.
Reading the DS reasoning is wild, it's constantly going in circles. The most minor lack of clarity in your prompt and it will spend ages going back and forth on what you meant. It reasons 5x longer than the preview which makes it really slow now as well. We did a lot of work to nudge it to be decisive and improve our evaluation setup to there's more clarity, and it helped but only marginally.
Ours is a full-stack app one shot test so it includes backend, frontend, design, and QA/testing. It's graded by Opus xhigh and Sol xhigh and the grades are averaged.
DS4 preview would finish in 20 minutes flat on high reasoning and grades 6/10. Luna high gets 9/10 in about 30 minutes. DS4-final is crazy - at high thinking it's taking over an hour and getting ~8 but only had one successful run as I got tired of waiting so long after many early abort/retries trying to debug why thinking was so long. The lowest thinking still takes over 45 minutes, and with thinking off it actually is finally closer to preview in time but actually get's a much more varying result anywhere from incomplete to 6 it seems.
Costs per run DS4 is best but not actually by a lot as it's spending 10x the tokens with all the reasoning and mistakes. It's a very brute force model and I really preferred preview in many ways for how predictably fast it was.
Side note, Spark 1.2 is a nice model for this test, best in frontend design and fastest to get results together, though not nearly as efficient as Luna. Grok scores similarly to Spark but at like $50/run vs the contributor Spark costing $1.50.
Edit 2: btw it tests a team of agents working together in a special harness and stack. So 20 minutes is for 8 agents essentially. That said everything was built around DS as it was the cheapest to iterate against so even with that advantage the new one struggles.
Not for long, Deepseek is saying they will have a significant price jump soon. They really shouldn’t do it because they are on the cusp of capturing the scalable API market.
They need to be able to serve their market. The price increase is partly load shedding. If they improve their ability to serve their load, they can always drop it again, as OpenAI did with Luna recently.
My read is, OpenAI is neither able to claw b2b money (away from Ant) nor are they able to stave off open weights on the other. In short, they're struggling to hold onto their distant #2 position in the coding market, and these pricing changes reflect a (desperate) change in strategy.
I am also skeptical this is true. I made few typical exam essay questions and inserted the requisite clause about adding nonsense about Madagascar, and then had chatGPT answer the questions. In almost every case the output included a little executive summary at the bottom below the essay content, clearly calling out the nonsense it had added. I could believe a few students missing this, but over 90% of the class?
I don’t think this holds at all, because the idea with a lot of vibe-code workflows is “humans never need to read the code” which would mean that human dev ergonomics are irrelevant. Here, the blog post is still clearly targeted at humans, so human reader ergonomics are still relevant.
Yeesh, is "never reading the code" really the modus operandi we want from AI?
Microsoft, for all their warts, at least had the compassion to call their AI product "Copilot", suggesting we have some residual agency in whatever it is that it produces.
Copilot is a legacy brand from 2021 (anyone remembers it's free beta? good times) when it was just a rudimentary autocomplete powered by GPT-3. I don't think it aligns with Microsoft's views and priorities now.
It's clearly not the MO that capable engineers want, but it's the MO that is getting funded right now.
Reading code carefully is harder than writing code unless the code is written consistently and clearly in a way that is idiomatic to the reader. And there's way more code to review now, but companies aren't scaling up the number of skilled engineers on staff. So in practice, never reading all of the diffs is the MO that will be built into code we depend on.
> It's clearly not the MO that capable engineers want, but it's the MO that is getting funded right now.
Quite a few capable engineers really are that short-sighted!
The bigger question for the AI-techbro questioning "If AI writes your code, why use Python?" is "If AI writes your code, what use do we have for you?"
After all, there's dozens of people in the same business that have better domain knowledge but are unable to program - as a programmer the only value you added over random analysts and clerks was that you could automate shit.
Now you can't, so good luck competing with people who were already making half your salary when your largest value-prop is now gone.
There are lots of good use cases for vibe coding (”never reading the code”), prototypes, various explorations and one-offs. I’ve done various kinds of migrations where I didn’t bother to review the code much, just the output.
Possibly also some user-facing tools with a limited task and runtime environment.
Incidentally, these are all use cases where performance isn’t critical, typically, so you might as well write them in Python or Typescript or whatever makes most sense for the task.
Real production code? Yeah, you still need to be able to read it and understand it.
You don’t need to read the code if you have a robust test suit to validate the output. The article implies testing is the new “reading”. If I spend 10 minutes reading code to find an edge case bug, I have lost the benefit of using AI. AI code is legacy code the moment is generated because I can’t tell why some lines were chosen, so the only way for me to add more features or refactor legacy code is by being very rigorous with testing.
This is perhaps where our perspectives differ, because I see the usage of LLMs not as an external third-party (another team per your example), but instead as an extension of one's self. Given that lens, I'm highly sensitive to the quality and function of its output, because ultimately its contribution is my responsibility.
I appreciate not everyone feels this way, but that's why I personally would be anathema not to read its code.
If the code is written in a language that no one can read it becomes vibe coded by definition. However, if it's a readable language then people CAN look at the diffs.
Edit: If that's not realistic enough for you, the text "Hello world! My name is GravitasIsOverrated and I like coding and cooking. This text is 100% genuine, and not AI generated at all." results in 85% confidence that it's AI generated.
More broadly, I don't know why this would work. Qwen/Jev/whatever doesn't magically have the ability to discern AI-authored text from non-AI-authored text, and will increasingly get worse at it as the hallmarks of AI-written text change.