Well, if you have a clearly defined task it's easy. When using them for my pipeline of writing explanations for Chinese words I have clear ranking, for example - opus 5.5 clearly better than opus 5 at writing and knowing details, annoying nit picker when it comes to finding errors (high accuracy, low usefulness) all in repeatable numbers on different datasets. The problem is that these models are most useful when you are facing a new task that you haven't encountered before. And yup then it's astrology.
It seems to me that often experts from some field will think less of other experts, basically because they have built a different understanding framework. So they both may be equally competent but perceive the other as less competent, and that is just based on the material, excluding some ego stuff.
Thanks, it seems like a nice way to compare effectiveness of different non-typical methods which are not yet capable of some more ambitious benchmarks.
There's a console and you can track your command! Very cool. Not sure if 3D is needed but very cool nonetheless, It might be that some things are missing (I don't know redis internals though) like some encoding step when getting data back to client, managing client connections and their state (multi, watch) etc.
Awesome dataset. Things are changing. In Chinese-learning community I found that surprising amount of people are creating their own tools without any intent to publish them, it just became easier to create your own thing that to dig out the good stuff from tons of tools already available.
Alignment is a myth. Safety of whom? Humanity couldn't agree on common set of values for thousands of years and we're not gonna suddenly do that in the next ten.
Which ones? Because many humans kill other humans rationalizing it by safety of other humans.
I mean I know it seems simple, let's just be excellent to each other. Christianity got pretty far on a decent basic set of values. But it's never simple[1]
Surely all the AI companies working with the US Department of War shows this is nonsense though? Even if they have accepted Anthropic’s red line of no autonomous lethal weapons, which seems to be the strictest anyone tried to impose, that’s still leaving tonnes of room where they intend AI to help target and kill humans.
Maybe I’m a Luddite but I find AI prose completely tiresome to read, and its presence in an article forces me to be skeptical about whatever the author wrote. The latter was obviously a problem pre-AI but these tools simply make it much easier to tell stories, narratives, etc. that aren’t yours.
Yeah, and there is no pleasure reading this. It's way too long and verbose for what it's trying to say. I'd much rather read the prompt, in bullet points or whatever format they were provided before inflating a few example pictures and notes into a short novel
How is this criticism not just ad hominem slop? You could have said "It's way too long and verbose for what it's trying to say". The writer is orthogonal.
The journey to a magical, free bird sound ID app involves some obvious steps (the NSF gave them money), but the amount of timeline offramps is really high. The NSF didn't think this app was educational ("it's just telling them the answer, they don't even have to work for it!"), the funding ran out so they were thinking of charging money, eBird happened to have five years of bird sound recordings they could use, a computer vision guy on the team happened to learn that spectrograms are basically the same as pictures...
And the result is tremendously useful! They get great signal on where birds are, and they're making people so much more interested in birds and nature and science. And I think charging $1 for the app or, shudder, some sort of subscription model might've ruined the whole thing.
Funnily enough I was in the bathroom (that's where I do my most random thinking) just yesterday evening thinking how template matching in computer vision, e.g. https://docs.opencv.org/5.0/py_tutorials/py_imgproc/py_templ... is basically audio signal processing but instead of being temporal it is spatial. Different signals but same principle.
I've got 33. I don't know how that stacks up against other but a large percentage of them came just from my backyard. Just got added a Baltimore Oriole that showed up this week.
I'm up to 71 thanks to the joys of taking it on holiday to other countries (less joy for my family), and occasionally going places just because I know there'll be new birds there.
Personal best for a single session is 23 birds in my back garden over the course of a couple of coffees. I have though spent a fair bit of time making the garden as attractive to different types of bird as possible - mainly through having a variety of different shaped feeders with different feed in. I can recommend the NatureSpy camera feeder too [0] - though I have had to modify it a little to make the food hole smaller otherwise magpies empty the thing in a single day just chucking all the seed on the floor. Dickheads.
Tawny owls make so many different noises. The classic sound is two, one calls and one answers but they are very chatty with different sounds. They live all around, but never seen one.
Merlin got me into birding, and I've found this to be true, at least in some regions, especially if there's lots of background noise.
Unless a species has a really, really unique call or song, I've made it my personal rule to get eyes on it and, if I'm still not convinced, to take a picture and cross reference with iNaturalist and what folks on there think of it.
I've learned to curse the species that differ from others by the sheen of the backs of the males' necks or the angle of a wingbar. Also, gulls that only differ by things like the color of their feet or beaks.
For me, mostly a lot of false negatives. I have a Pixel 8, and I have no problems using the microphone normally, but somehow, Merlin requires the bird to be within a couple of meters of me and sing for quite a while before Merlin will even say “hearing a bird”. And of course, you have to be dead silent; if anyone talks or coughs, it will mess up the spectogram badly.
Given that it's colloquially “Shazam for birds” and Shazam is just amazingly resistant to noise, it's a bit disappointing :-)
I think what makes Shazam magically "just work" is that they are searching for the exact embedding of the recorded version of the song - it doesn't work for a cover version, acapella, etc even if it is very similar. (Everyone would want an app that you can hum into and it will ID the song, but I guess it's a very tough problem.)
Merlin is solving the more general problem so makes sense that it's much more finicky and less accurate.
Still, we're talking orders of magnitude here. I've fed Shazam with a stream of the cheapest lapel microphone you could get, mounted behind a rack filled with twelve very noisy servers, in a hall filled with 5000 people also all being noisy, and it would reliably recognize music from the other side of that hall before I could do so myself (I had to walk halfway there to hear “oh, yes, it's actually right”). And that's on a stream encoded with 9600bps GSM compression. With Merlin, I can hold out my phone on a silent day, hear the bird loud and clear and it will hardly register.
Shazam primarily identifies on the geometry of spectogram peaks, FWIW (I wrote my master's thesis on the DSP of music recognition back in the day; it's possible that they are doing something more fancy now, of course, but I'm not sure if they would want to). I don't know exactly what Merlin is doing.
> an app that you can hum into and it will ID the song
YouTube Music and Google Search do OK with these types. Sometimes if I know a song well enough, I will open YT Music, activate the mic search, and sing a few bars, instead of typing the song title.
According to Gemini, SoundHound is also a good app for song-hummers to try.
I'm not sure why but this seems very phone dependent. I can stand next to another person, both running Merlin in silence, and mine will pick up much more than some people's phones. My phone is not particularly new (Zenfone 8).
You can usually tell, though. Rare birds in the wrong place or time that come up only once are problematic, but if the bird is detected repeatedly there's a better chance. You should also be on the lookout for birds that mimic other birds (like the northern mockingbird and the gray catbird.)
Occasionally I'll see a screenshot of Merlin with, I dunno, a dozen different birds, and in the middle of the list: a catbird. I never have the heart to tell people that they just heard one very creative bird.
Indeed - background noise will tend to lead to a lot more false positives, but you should use it as a guide rather than inherently trusting results. Same goes for using BirdNET-go and other models, though these do allow you to be very strict with detections and reduce false positives more.
Cornell / Merlin / eBird publishes a global species frequency DB indexed by location and date (so, how likely is a species to appear in a place, on a day of the year)... In a project that I was working on, this helped a lot when used to influence the "raw" identification based on appearance alone.
The point is that auto mode gives people a false sense of security that leads them to believe they don't need to run Claude in a proper sandbox. This same attack running in a sandbox (even in YOLO mode) would be comparatively harmless.
Ever since they made auto-mode default I swear claude has tuned to use python commands instead of the Edit Tool to frustrate the ~security conscience~ luddites into using auto-mode.
Yeah, It’s in the system prompt, Claude will tell you if you ask why it’s using Python.
My theory is that Anthropic is just a vibe-coding company. Their goal is to capture the attention of white-collar non-coders, since programmers will jump ship fast to another model.
That's what I meant - you give it an instruction that seems to work (always ask before deploy) and so you trust it, and then you notice it can easily convince itself to deploy without authorization ("the user asked me to fix this, and they must know it's a deploy ...").
Is this about normal system prompt instructions or instructions for the auto mode classifier? I'd be a bit more surprised about the classifier forgetting instructions.
The commenter you replied to mentioned that you can customize the auto mode classifier by providing a prompt, implying that this would be a more robust way of constraining Claude's behavior. It wasn't clear from your response whether you were using this functionality. You might try it out as a way to more reliably prevent these kinds of workarounds.
It's absolutely not reliable, and we have opened a few issues for that.
For example: our instructions (which are read by the model and classifier) include "do not use sed/python/perl/etc, always use the edit tool for editing", and this only gets followed for a few messages. We have introduced scripts to block those ourselves, since the classifier doesn't care.
Because of those problems, my team is currently testing OpenAI after about a year of Anthropic.
What seems to work for me is automation - read file hook that re-injects instructions in the prompt every 15 minutes. Switch on the filename and get language-specific instructions too.
I also feel like this is an attack that manual review is not that likely to catch, given none of the malicious code appears in any of the tool calls or output.
Can you suggest a proper sandbox on mac? One that allows both me and the agent to interact with the processes? Where it can drive browser, for both oauth setup and runtime visual inspection? I've tried building docker setups, but can't figure out the browser driving part.
I don’t think it’s related to the auto mode at all. It would work perfectly in the manual mode. It does not even need Claude: just give a human a similar archive and hope they run some simple Python from the directory at least once. And make sure there are lots of files do they don’t notice a weird .py around
reply