I'm getting about 3k tok/s on prefill so it heavily depends on the input size. I hooked it up into a coding agent for shell permission checks, estimated shell execution time (buckets), goal pre-screening, subagent routing etc. and I have RTTs between 100-200ms. It can take noticeably longer for huge contexts because of the prefill speed but it can be cached so subsequent calls will be much faster.
I tried smaller decision models like laya (also with custom finetunes) but the accuracy was not really good (for the things I tested). Also I don't have any VRAM left on this GPU so i had to decide whether to host laya or qwen-3.8-27b but not both at the same time. Running decision models on a CPU will also be noticeably slower so I went down this route to have both combined with shared base weights.
we use that same model at our company, it powers not only our devs but also many business needs. We rent 1 gpu (B300), it costs around 10x less than the api costs
Do you rent from AWS or some other provider? I wonder if there are issues with on-demand rent for those high-end GPU instances, is capacity always there or sometimes it is unavailable?
If by "genetic modification" you mean "random genetic mutation" formally known as stenospermocarpy.
Seedless grapes existed in the Mediterranean centuries ago. The now-popular Thompson seedless was brought from the Ottoman empire in 1872.
Similar story with bananas and the parthenocarpy mutation. Then you layer selective breeding on top of that to create complete sterility due to triploidy.
Others like the babaco papaya are just simple sterile interspecific hybrids.
So nature actually does make seedless fruit by accident quite often. It's just that you need humans to come along and clonally propagate them.
wow, what kind of stuff does that script do? I've seen non-vision models analyze images, but mostly histograms, color averages etc. This one seems to actually understand the image itself and reproduce the layout, very impressive
reply