I wanted something that I could easily configure to manage a set of sensible defaults, that supports multiple llama-server binaries, with per-model over-rides, and command line over-rides. The utility is here: https://github.com/stew675/start-llama I know that llama-server has its own model loading