inferencecontrol-planenvidia
One ModelDeployment, two serving stacks
A Modelplane cluster can now serve models with NVIDIA's Dynamo components instead of the stack Modelplane composes itself. It's a per-cluster platform choice, and the API an ML team writes doesn't move.