@billlycoder: Everyone is posting this as open source. The code is. The weights are not. NVIDIA's grounding model predicts a whole bounding box in one step instead of spelling the coordinates out token by token. 12.7 boxes a second on one H100, against 1.1 for Qwen3-VL. Code is Apache 2.0. Weights are research or evaluation only. Comment BOXES and I will send you the repo and the licence. #opensource #nvidia #computervision #ai #coding #devtools