Share a GPU between local models
Several services on this server use the same 32 GB graphics card. Dictation needs about 2 GB in the recorded setup, the larger language model uses over 20 GB, and image generation can need around 20 GB. They cannot all keep that much memory at once.
These notes describe the GPU deployment. It has not been repeated in the guide-testing VM, which has no NVIDIA card.
Check what is using memory
Before changing model settings, look at the processes already using the card:
nvidia-smi --query-gpu=memory.used,memory.free --format=csv,noheader
A service may still hold a model after finishing a request. In this setup, releasing an idle image-generation model made room for the next job.
Decide what stays loaded
Dictation stays available because a loading delay interrupts its use. The larger chatbot model can load when needed. Image generation runs when enough memory can be made available.
Your priorities may differ. Note each service’s memory use, loading time and whether it is acceptable to interrupt or unload it.
Coordinate model loading
A helper in front of the model server checks free graphics memory before admitting a large request. It releases eligible idle models, waits for memory to become available and then lets the request continue. Protected services, including dictation, are excluded from unloading.
The helper must also deal with concurrent requests. Checking free memory is not enough if two jobs can both pass that check and start loading together.
Work within the remaining limits
This arrangement lets the card serve several jobs at different times. It does not make their combined memory requirements smaller. A model may still be too large, or a request may need to wait while another job finishes.
Measure the models and workloads you actually use before deciding whether scheduling is enough or more hardware is needed.