The process by which benchmarks are setup and run does not correspond at all to how human developers engage with a coding agent. At best it is a loose proxy, and often a bad one.
What benchmarks are usually good at is showing to what degree new models are better than old models. What they are not good at, by construction, is showing that harnesses are well adapted to how people use them.
I presume you're not talking about autonomous agents, right? Because then you could give it a set of tasks and check its success rate. But, even so, are you not giving it tasks? Are you principally interacting with it through discussions that are more difficult to quantify? Even question answering has benchmarks. I'm having trouble imagining how you use them (or "how people use them" in your words).
I don't know if you noticed, but Opus 4.6 was peak for human-computer interaction. Everything has been fairly downhill from there despite better benchmarks, at least in that one regard. Opus 5 and 5.5 are clearly a step above in capabilities than 4.6, and I don't think anyone wants to go back, but 4.7 and 4.8 were arguably worse overall. I genuinely feel I got more done with 4.6 and often switched back, prior to 5 coming out.
Why? Because 4.6 actually talked like a human being. It actually organized its thoughts well, and got the main information across without the wall of text that makes your eyes glaze over. So from the perspective of human-computer interaction and maximizing the productivity of a developer+agent team, 4.7 and 4.8 were regressions. Despite much better benchmark performance.
Even if we consider autonomous agents, that benchmark is not indicative of how well they will interpret *your* requests. Or how well they will interact with other agents in a flock/swarm situation. The benchmark just doesn't cover this. (And the difference can be nontrivial! Sakana AI's published results show two generations of uplifting potential from better harnesses.)
They could, but as you see here, people are very eager to create dashboards and trackers that do external accounting by proxy, so they can't just "make up numbers" without the customers noticing and making a fuss.
What do the engines have to do with plasma around the hull, or rather its leading edges, caused by supersonic speed, which never ever even enter them?
And if you're thinking of the exhaust, why would that matter at all for signals reception and transmission, with the exception of a small cone from right behind?
It's not something I can cite with any public URL to a website that's an authoritative source, but if by 'spy' you include SIGINT satellites in that category, yes they absolutely do... Ever seen a diagram of a Boeing (Hughes) 702 series satellite bus and antenna system built for L/S-bands like frequency ranges similar to those used by Thuraya, Inmarsat and Iridium handhelds?
It's widely accepted in the commercial two way satellite telecom industry that the NSA has comparable things with big unfolding umbrella-like antennas in geostationary which serve a similar purpose for signals collection.
> It's widely accepted in the commercial two way satellite telecom industry that the NSA has comparable things with big unfolding umbrella-like antennas in geostationary which serve a similar purpose for signals collection.
How does that even work, do people just accept that their satellites hang around suspiciously close to other geostationary satellites and hoover up the very edge of the beam somehow?
Because if you're transmitting stuff *to* geosynchronous birds you have to aim pretty accurately.
Handheld S/L-band things like Thuraya, Iridium, Globalstar, Inmarsat phones don't transmit in a very directional way at all. Or things like a Garmin inreach that have an Iridium modem in them. They're omnidirectional enough that the broad requirement for a user is to be outdoors with a good view of the sky.
If you have a NSA-type geostationary satellite somewhere in approximately a 1/5th arc of azimuth along the equator of the earth that can 'see' an area on the ground, with antenna systems aimed at the area of interest (let's say, Afghanistan), it's not necessary for your NSA satellite to be anywhere close to the actual service provider's satellite.
Note that Thuraya BGAN-like devices and handheld phones talk to geostationary, Iridium is LEO (but likely some combination of LEO, molniya orbit and geostationary can "sniff" it), Globalstar is similarly LEO in its service provider architecture, and Inmarsat is geostationary.
Also lots of conjecture out there that sigint platforms exist which are good enough to capture ordinary handheld device cellular traffic or VHF/UHF radio traffic that's not intended by its users or terrestrial operators to go into space at all. This is probably done from a combination of low earth orbit platforms and more complicated orbits or what we would call MEO (maybe about the same height as an o3b satellite?) to achieve longer dwell times over an area of interest.
But then with the difference in distance from LEO stuff like Iridium (RIP - I actually saw the ground flash of an Iridium flare once!) you'd need a truly *massive* antenna to hear that out at geosync.
An Iridium 9555 handheld phone talking to Iridium satellites in LEO with its antenna extended transmits in approximately the same frequency range and about the same tx power as an inmarsat isatphone 2. The former is LEO, the latter talks to geostationary.
Same concept for an Iridium 9555 or its successor in a dock plugged into the small coaxial cable that goes to the external hockey puck sized antenna. This is all happening between approx. 1400 to 1700 MHz in either system.
Look up the detailed RF specs for the isatphone 2.
Both of these only work because the channel size is extremely narrow.
The antenna on the commercial satellites that do two way S/L-band to handheld stuff (or INMARSAT BGAN size terminals) from geostationary is massive, I've no reason to believe that a high budget billion dollar NSA satellite at MEO or higher altitudes wouldn't have a similarly gargantuan unfolding antenna.
The ground units do not have any sort of narrow beam capability, thus you don't need to be anywhere near the satellite to intercept them. I have a low end satellite device: Garmin InReach (text only, provides tracking and emergency communications in the middle of nowhere.) It has a preferred orientation, not a required orientation.
Why do you think I am confused? Is this a form of internet condenscension, or are you able to diagnose my mental state based a single comment? No, I don't think I'm confused.
Because GEO is about 100x further than LEO, which means the signal is about 10,000x weaker. Also the risk to geo communications from having multiple birds in the same slot is huge and treaty violating. Why put that risk and insane demand when you can still get global coverage from a dozen or so birds in high polar orbits?
Nobody does what you are saying, because physics, which leads me to believe that you are confused.
If you think the US doesn't have a plethora of DoD things in geostationary orbit you would be extremely, terribly wrong. In addition to well known programs like AEHF, WGS for two way military satcom purposes, a person with an amateur grade telescope can literally spot the satellite that keeps station near an existing commercial Thuraya satellite:
reply