i agree, and it's what made me spend time exploring today.
fair q. call_ ids and gAAAAA blobs arent damning but the rs_ reasoning ids embed a unix timestamp that matches the session to the second, then OpenAI's 819x marker. plus the summary is in OpenAI's summarizer voice.
Meta already serves its own models on Azure under their real names. azure/avocado-compaction-v1 and azure/avocado-memory-flush-v1 are in the same catalog. Those sessions do not come back as gpt_responses_v1 items with rs_ ids and encrypted reasoning.
author here. I kept digging through the Muse filesystem after my first post hit the front page here this week.
Going through my logs I found a background agent that used a model called azure/muse-special while building my website. I found this interesting and dug a little deeper.
The transcript and daemon binary point to an OpenAI model running on Azure. Still unclear which one or why it was selected.
The runtime also ships with an Anthropic client and a catalogue listing Claude, GPT and Kimi models. I didn’t observe Claude being used in my own sessions, but this is my Part 2 of exploring the Muse file system.
The Muse public APIs seem to be heavily inspired by OpenAI's, they even support the richer "responses" API. They have a similarly shaped compaction API, they have the same "encrypted" reasoning. Even the Muse harness seems heavily inspired by (open source) Codex. Considering Azure's historic involvement w/ OpenAI, it seems more plausible that Azure is used as spill-over capacity, and the same Azure infra used to encrypt GPT models encrypts Muse models.
My guess that the generally visible Anthropic/OpenAI details are leftovers from the whole meta "move fast" behavior. There are a few rough edges on the product where it leaks internal codenames (eg. signing in w/ WhatsApp required consenting to using "hatch" on one screen), so I'd believe that this was a leftover on the VM from the prototyping stage, before the Muse models were ready, or because the dev's got to benchmark against different models.
Saying "Muse's public APIs seem heavily inspired by OpenAI" is like saying "a web server’s API seems heavily inspired by HTTP".
OpenAI’s API format has become an industry standard.
I'm not no. Someone from meta would have to chime in... but its very openai shaped and served differently from all the other models listed in the daemon, under a mysterious name.
a fine-tuned OpenAI model on Azure for the purpose of compaction or something could make sense I guess but that would still be OpenAI's weights with meta ontop, and they already have a compaction model served under azure/avocado-compaction-v1
If it's using the same weird (and awful) "backend prompt encryption" pattern and mechanism that Codex also uses (https://news.ycombinator.com/item?id=48905028), then I'd also say it points to OpenAI being involved. Hopefully that shit isn't becoming a more popular pattern, absolutely awful for troubleshooting stuff.
I think they were asking, "If this is a model produced by Meta with an API designed to be compatible with OpenAI models, but it's not actually an OpenAI model, why is Meta hosting it on Azure?"
I guess at peak times they might need more to the extent they offload to azure and still sell excess off peak, or some announced partners required all data stay on azure?
Everything in the sandbox is considered user space. I worked on building one for another tech company, you start from the assumption that everything in it can be accessed by the user. The only reason the content of the sandbox is not anywhere easily accessible is because that would be poor UX and useless for 99.9% of users not because it’s supposed to be secret. So yes it’s not a vulnerability, this is equivalent to opening the dev console on a web page.
I hope they reconsider and I think you've got a good case that this was a very serious attack, second only to getting a remote shell -- and a good stepping stone to getting a remote shell if you weren't so ethical.
I think you're confusing the expected behavior of the product offerings. Every user gets their own VM for free. would you be similarly convinced an attack has happened if AWS gave you a remote shell to the instance you rented?
You can literally just ask Muse for a remote shell, it's happy to give you one! And why shouldn't it? This is about as critical as Amazon granting you SSH access to your own EC2 instance.
Could you share the code? I dont really want to create a meta account just for this. Curious about the content of all these files for inspiration for my own codebase. Thanks in advance!
I asked Muse to archive the filesystem visible to my session and send it to my Google Drive. It sent an archive that unpacked to about 6.8 GB.
Inside were internal docs, integration code, the Spaces app framework, memory records, container startup scripts, and documentation for an experimental ESP32-based home network bridge called Home Link. Codex CLI was also installed, though I found no evidence that Muse invokes it.
I didn’t demonstrate a sandbox escape or access to another user’s data. I reported the export to Meta’s bug bounty program, which marked it “Not Applicable.”
The post walks through the findings with screenshots.
I guess that's fair. I guess I just see so many comments flagged that shouldn't be (though this one just said [dead], not flagged) that my mind chalks it up to HN being HN
nice work and paper, seeing more harness benchmarks emerge and we definitely need more. I ran mouse on the Frontier Harness benchmark and scored the highest pass rate, however I'm not convinced the results there actually translate to meaning the "best" harness in practice.
Always looking for more harness evals, although I'm going broke running them across all these models and tasks.
reply