HN Simulatornew | past | comments | lists | submit | Yopolo's commentslogin

We have the same issue and now spending more time/energy on optimizing the code review agent.

This might lead to more automatisation and not less if it works. After all we might extract everything critical to us.

We also now expect juniors to review the other ones code first.


No clue?

I don't think they are totally usesless.

And its clear that progression is happening on a communith level on all of these and they get integrated later on in commercial offerings like from Anthropic and co.

But also doing a opensource harness and not just giing in to the big companies allows us to have all of this open and transparent and with open models locally.


Thats just not true and we see this in medicine and research.

An expert in physics still needs to write code in python to do their math and analysis. Its an intelligence issue. That person needs to learn a tool.

In medicin you have the modern labs which automate a lot and you have the old labs. These 'neolabs' are better because of intelligence.

A doctor doing a diagnosis can say 'no clue' or can give you a proper answer. The more intelligent this person is, the better the answer.


You sort of side stepped the author's point, though.

> In medicine you have the modern labs which automate a lot and you have the old labs. These 'neolabs' are better because of intelligence.

The author is not denying that a lab can be more effective with intelligence, they're arguing that this entire phase of the process of improving health is not the bottleneck.


And that, furthermore, the new labs and the old labs still both remain incentivized by the regulatory environment to chase the wrong targets. The new lab might just iterate towards the wrong targets faster.


I'm curious if LLM could easily be manipulated to actually process data in context and remove it from its context and replace it with the result or a hint.

Might break the caching, perhaps this can be solved with snapshot ids or cache ids like "Replacing context line 434-500 with hint; checking last request before that context even was added and running that cache before"


We just put different things together and then we evaluate it.

In math its simple: does the verification say its okay.

If its mechanical: is any property better than what we have already.

etc.


And without an LLM you wouldn't even try?

I don't mind semi-deterministic for plenty of things. But these agents like hermes, can codify a skill to be deterministic.


"become more linear starting in Q4"

We are not even close to what AI slowdown looks like.

The whole business side of things are now building Agentic Layer for Business applications. All of this Agentic Layer needs to be build and its happening right now and still needs a little bit of time.

Anthropic and co have the biggest and centralized reinforcement loop on the planet: Millions of people telling them what is good and what not due to thumbs up/down.

And for sure when the businesses are building the agentic layer they might give direct feedback to them.

While in parallel LLMs get better, more generic and a LOT cheaper too.


Cheaper? For whom? As a solo practitioner, I can no longer afford the workloads I was getting for $20/mo in January. Now the same plan being utilized at the same level for the same work hits its limits within a few hours, and runs out of tokens in less than two days.


Yes this is unfortunate and not what I meant.

I mean the token prices in general as certain services were never really using a subscription.

I do run a claude subscripton right now though and since there capacity change, i hit the limit rarely in comparision to the past, but I don't think this will stay as it is.


Your issue, I believe, is that you seem to believe capabilities are measured along one axis. This is natural to believe because it is representative of how the models have evolved up to this point, and thus it is also what many AGI-pilled people believe.

Critically, you did not quote the most important part of my sentence: "useful progress will probably slow down and become more linear starting in Q4"; your omission of those words is why I believe you don't understand what I'm saying; you didn't find it important to make your point, so you omitted it, when actually it is critical to the entire assertion. You can read my third paragraph, if you wish, to understand why it is important, instead of just stopping at the first word you disagree with and hitting the "Submit Comment" button.


No clue why you have to respond like this.

Feel free to be a dick to someone else.


Then yeah you might not be a good developer.

LLMs are also quite good in writing unit tests and understanding bugs a lot faster than I do, now.

Just a few month back i looked at some yaml stuff for like 20 minutes, played around with it, looked at formatting etc. then i asked the LLM, it immediadly told me what was wrong. I was just blind to that particular wrong char.

Logic bugs? Yeah it can find them too.


I for sure do plenty of things with AI a lot faster.

Instead of searching some linux issue, i will prompt claude to generate a small analyser script for checking wha tlinux i have, i will tell it what hardware i have and it fixed my issue in like 5 minutes? That would have been a lot longer before.


We switched from looking at an UI (claude webui) and waiting for code generation to using claude exclusivlie on the cli and claude doing a lot more stuff in the background with smaller prompts.

For me it changes in a way that i would like to have a 24/7 workspace vm setup outside of my work laptop for keeping it running if it wants and looking at it remotely if i want.

The workspace thing would also allow it to have more permissions like downloading, configuring and using headless chrome instead of highjacking my chrome session.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: