I think there will be a lot to be learned if we can scale up the high resolution measurement methods here to map exactly how brainwaves travel when performing specific tasks.
Maybe they should recruit experienced meditators who can give accurate introspective self-reports to be compared to the measurements, then we would start to have a relatively confident map between brain physiology and psychology. (e.g. X brain pattern in Y region specifically corresponds to the Z step of reasoning when solving the problem)
I'm pretty convinced now that grading take-home assignments is pointless with the prevalence of AI cheating. (There's surely already "innovation" going on targeted at letting people cheat on assignments while staying undetectable.)
It's probably best going forward to grade only in-person proctored exams using paper, or offline air gapped computers in the case of programming courses.
I had many classes in college that didn't administer homework or take-home assignments. All we had were just exams. Most of those professors operated on the logic that if you wanted to pass their class, then you would do the (suggested) work.
Also, I never had a computer science class where exams were administered on a computer. I only graduated like 10 years ago, too. I shit you not, even my x86 assembly class was all pen and paper too, and yes, we were graded for code accuracy.
Afaik the usage quota nerf to $200 Pro is true. Twitter is all over them right now. (Also you haven't been able to buy the $200 Pro plan for the last two weeks or so)
They're trying really hard to not mention that the obvious practical solution to the lack of frontier models in cyberdefense is that the defenders should run GLM 5.3 themselves.
Perhaps they should do something like remove dual-use cyber safeguards on older models as soon as open weight models of a similar capability are released.
Lock-in is easy with these agents because they NEED all of your data to be useful, and they will continue to learn internally about you.
But ultimately the AI company CAN choose to just make all the data exportable and open source their product for self-hosting. (The mainstream ones won't, of course, they want to lock you in and hide their AI prompts and algorithms.)
I think what is sorely needed is a version of Dots/Muse without lock-in risk but is still accessible to regular people unlike Openclaw.
I don't think export is necessary. Most of the value is in the context, you can always prompt your agent to vomit all of its context out into markdown files in some repo and then pass it on to the next agent
i'm usually not one to defend europe's (over-)regulation, but having ownership of such data and being able to direct it to other providers for easy switching is one of the pillars of a lot of data related regulation
so either they can forget about their moat or forget about the continent with those practices
I agree. FWIW I'm trying to build a wordpress-like AGPL 3.0 thing (holdout.com) that lets you import your existing chats, has feature parity with the frontier consumer apps, rich plugin marketplace, and hopefully lets people feel more in-control, and able to sacrifice less of their entire lives to one walled garden.
Yep, that's what I think and wrote on my blog. My hope is that the inference providers should provide a managed service like it.
It's also in their best interest to do so, because if they don't and these provider hosted agents become the norm, demand for independent inference would drop.
> It's also in their best interest to do so, because if they don't and these provider hosted agents become the norm, demand for independent inference would drop.
I disagree with this part. AI companies will want to try their damndest to control distribution of AI, so that they can enshittify later.
Consumers conscious of this will want an alternative, of course. Might be niche similar to how Kagi is in search because big tech will always have a AI inference cost advantage + making users the product (extra $ from ads & purchase cuts) + the good old strategy of dumping.
Very similar in principle, but I'm aiming for something very different in UX - namely, a polished client app built for creating and messaging with an agent that coordinates work across many subagents
I find Telegram, Messages and the like to lack the fidelity I'd want for managing such stuff. And the official OpenClaw app is ... well, it's weird that they shipped that
> inb4 "humans aren't reliable either." You, person typing that very response
> out right now, you are part of the problem.
The problem being that we are adults, have to deal with real life, and can no longer pretend that running to mommy when we don't understand something gives us a "reliable" answer? You may next find out there is no Santa Claus.
This lazy style of participation sounds like “wow this article touches on something I happen to have an opinion on, now is a good time to share it.” You’re reading headlines and not engaging with the material present.
> indie game developers have stopped the traditional practice of updating a dev log as a form of marketing and community building because people will point their LLm at it and create your own game before you do
And that's why we can't have good things...
(This is just gonna keep happening more and more until eventually we'll need something like a patent system for ideas)
hmm - those tools are great, but I don't think orchestrator / ADE is quite right either, since this is explicitly not opinionated about where your coding agent is living.
Whiteboard is targeted at almost the opposite problem of the ADE (which is targeted to context switching) - having a dedicated tool to help you understand & participate in the development process in places where humans are high leverage.
This is definitely getting at least some things right about how we work with agents today, specifically that we often work at the architecture level, and we need a better alternative to the current Plan Mode offered by coding agents to efficiently architect software at a high level, which is more visual and offers better back-and-forth incrementation with the agent than simply "reject final plan with X message".
From the website demos i definitely think this is a clean interface, although I don't know how much better this is compared to some simple custom Mermaid format, which the agent can write as artifact files and present to users. Zooming out, this app seems like 1 feature (a MCP with a GUI attached to it) rather than an entire product.
Also, I don't know if asking the agent to write specific code changes into the plan is a good idea. I think maybe that a "plan -> approve -> write code" would let the agent write higher quality code than "plan which contains code -> approve". But maybe you can make it work when combined with some specific prompting marking the code as clearly work-in-progress and subject to change, and that the agent should surface any parts implemented differently relative to the plan to the user, etc.
1. "is this a feature" - it could be! in fact, we will expose this as an MCP UI next so that you can view the info directly in Codex Desktop or Superset/Conductor/Emdash for example. that aside, we found that the big things that matter for us are: (1) good code navigation (diagram/spec -> code), (2) beautiful diff viewing, and (3) visualizing agent traces as they connect to code. we found that these problems were hard enough, and enough folks that were using platforms that didn't easily map to these requirements - e.g. TUIs like claude code - that a dedicated product that was just focused on these problems exclusively makes sense.
2. "using whiteboard for plan mode": hmm, i think our wires are crossed a bit here. how people mostly use whiteboard today is:
plan -> approve -> agent codes -> use whiteboard to explain the code.
(or just omit the plan phase as a formal artifact -> just emit a plan + code together, like a golang design draft [A]).
we are exploring an explicit "put the plan in whiteboard first" mode (there's a scratchpad feature that's experimental right now), but it's definitely not ready for prime time yet.
> how people mostly use whiteboard today is:
plan -> approve -> agent codes -> use whiteboard to explain the code.
I guess it's interesting and useful for now, but I don't think people are going to work at the code level much longer.
In my opinion current coding agents + automatic review systems are already at superhuman reliability during the implementation phase (as in they will not fail something in the plan during implementation and not tell you about it, so there's no need to look at the actual code beyond maybe a cursory glance). I literally just use plan mode + CC's /code-review in each task so it's not like I'm doing anything special. So I think the main human interaction surfaces to target in the future will be in the planning process.
Without going through the following flow I very quickly end up with nitty recurring micro bugs or interaction "gunk" and 10's of thousands of lines of pointless code that bloats out the code-base and agent's context.
plan -> approve -> agent codes -> review code -> cleanup and pointing to specific skills/agents depending on what the issues are and some manual instructions pointing to specific lines of code-> agent codes -> review code -> final cleanup -> ship.
Working on a 3d sims-like video game I cannot get away with much less than that flow outside of very small features.
I have (like all of us I assume) tried moving forward over weeks of not code reviewing and only plan reviewing, and it's amazing how badly things fell apart, I ended up having to reset weeks of work and ended up knocking out like 50K loc for the exact same features and zero bugs instead of constant bug (or just interaction/latency gunk cropping up everywhere) cleanup slowly escalating and agents needing to take hours to build anything.
> So I think the main human interaction surfaces to target in the future will be in the planning process.
yes, agreed. we're working on more stuff in that direction (a plan / scratchpad mode), but what i personally like the most is eliminating / shrinking the plan/review gap.
i think reviewing a plan without an implementation doesn't feel that useful anymore, at least to me, because key tradeoffs often only surface during implementation that effect the top-level spec.
in some sense, the code writing process is just a cheap effort which makes the spec better and more thorough?
Is that really true? I feel even with Fable and the likes, once the LLM has locked down an implementation any reworks I try to do gets it really tunnelvisioned on the current implementation, treating it like the truth even though it JUST wrote it, and any attempts to make it reframe the problem just makes it dig down further. In these cases I always get way better results throwing the whole thing away and rewinding the conversation rather than trying to evolve it.
An easy upgrade (ime) is to be intentional about a process, move the planning artifact to a file, use multiple research/propose/review sessions to dial it in. Still tuning my vibes for when to add in some actual exploratory implementation elements, because there's always something you didn't foresee when getting to the actual implementation, while also not having them implement the solution as a "plan" in markdown
> be intentional about a process, move the planning artifact to a file, use multiple research/propose/review sessions to dial it in
Yeah, I think we need something like that as well. I am actually working on an virtual artifact filesystem in my orchestrator to enable this. So agents can create a persistent, versioned plan artifact separate from the codebase (maybe a HTML) and iterate it alongside the user, much like what ChatGPT/claude.ai can already do but for a coding agent. Then you'd need to define a process and get the agent to follow it, but that's much easier and mostly a mix of prompt and orchestration primitives.
> exploratory implementation elements
This is a good point, I've ran into a lot of instances as well where my agents in plan mode would like to explore something but can't because of permissions. I wonder if there should be some kind of system like a "experiment subagent" to handle it.
Commit it to git, it's not far off from an llm-wiki
I have no orchestration primitives, just a skill tied to a .design/*.md
Unless we consider opencode sometimes using a subagent as a primitive? Maybe the problem is leaving the clankers to their own devices for too long/much
I intentionally block almost every tool for the design/review agents, letting them "do" things is a distraction. I will use the build agent and tell it what to do if I need that experiment. I don't want to have to read through the wasted tokens a bunch of dumb bots burned through to create walls of markdown. They go on way too many side quests
Maybe they should recruit experienced meditators who can give accurate introspective self-reports to be compared to the measurements, then we would start to have a relatively confident map between brain physiology and psychology. (e.g. X brain pattern in Y region specifically corresponds to the Z step of reasoning when solving the problem)
reply