I've not committed to that as a benchmark yet because you really need a full coding agent configured to run Blender, and I like benchmarks I can run as a single prompt/response through the appropriate API.
I have been intending to get more of a coding agent benchmark going though, so maybe this should be part of it.
I've not committed to that as a benchmark yet because you really need a full coding agent configured to run Blender, and I like benchmarks I can run as a single prompt/response through the appropriate API.
I have been intending to get more of a coding agent benchmark going though, so maybe this should be part of it.