And its clear that progression is happening on a communith level on all of these and they get integrated later on in commercial offerings like from Anthropic and co.
But also doing a opensource harness and not just giing in to the big companies allows us to have all of this open and transparent and with open models locally.
You sort of side stepped the author's point, though.
> In medicine you have the modern labs which automate a lot and you have the old labs. These 'neolabs' are better because of intelligence.
The author is not denying that a lab can be more effective with intelligence, they're arguing that this entire phase of the process of improving health is not the bottleneck.
And that, furthermore, the new labs and the old labs still both remain incentivized by the regulatory environment to chase the wrong targets. The new lab might just iterate towards the wrong targets faster.
I'm curious if LLM could easily be manipulated to actually process data in context and remove it from its context and replace it with the result or a hint.
Might break the caching, perhaps this can be solved with snapshot ids or cache ids like "Replacing context line 434-500 with hint; checking last request before that context even was added and running that cache before"
We are not even close to what AI slowdown looks like.
The whole business side of things are now building Agentic Layer for Business applications. All of this Agentic Layer needs to be build and its happening right now and still needs a little bit of time.
Anthropic and co have the biggest and centralized reinforcement loop on the planet: Millions of people telling them what is good and what not due to thumbs up/down.
And for sure when the businesses are building the agentic layer they might give direct feedback to them.
While in parallel LLMs get better, more generic and a LOT cheaper too.
Cheaper? For whom? As a solo practitioner, I can no longer afford the workloads I was getting for $20/mo in January. Now the same plan being utilized at the same level for the same work hits its limits within a few hours, and runs out of tokens in less than two days.
I mean the token prices in general as certain services were never really using a subscription.
I do run a claude subscripton right now though and since there capacity change, i hit the limit rarely in comparision to the past, but I don't think this will stay as it is.
Your issue, I believe, is that you seem to believe capabilities are measured along one axis. This is natural to believe because it is representative of how the models have evolved up to this point, and thus it is also what many AGI-pilled people believe.
Critically, you did not quote the most important part of my sentence: "useful progress will probably slow down and become more linear starting in Q4"; your omission of those words is why I believe you don't understand what I'm saying; you didn't find it important to make your point, so you omitted it, when actually it is critical to the entire assertion. You can read my third paragraph, if you wish, to understand why it is important, instead of just stopping at the first word you disagree with and hitting the "Submit Comment" button.
LLMs are also quite good in writing unit tests and understanding bugs a lot faster than I do, now.
Just a few month back i looked at some yaml stuff for like 20 minutes, played around with it, looked at formatting etc. then i asked the LLM, it immediadly told me what was wrong. I was just blind to that particular wrong char.
I for sure do plenty of things with AI a lot faster.
Instead of searching some linux issue, i will prompt claude to generate a small analyser script for checking wha tlinux i have, i will tell it what hardware i have and it fixed my issue in like 5 minutes? That would have been a lot longer before.
We switched from looking at an UI (claude webui) and waiting for code generation to using claude exclusivlie on the cli and claude doing a lot more stuff in the background with smaller prompts.
For me it changes in a way that i would like to have a 24/7 workspace vm setup outside of my work laptop for keeping it running if it wants and looking at it remotely if i want.
The workspace thing would also allow it to have more permissions like downloading, configuring and using headless chrome instead of highjacking my chrome session.
This might lead to more automatisation and not less if it works. After all we might extract everything critical to us.
We also now expect juniors to review the other ones code first.