whichllm: pick a local model that is actually fast
whichllm ranks the best local models for your hardware — general, coding, vision, math — then whichllm run downloads the top pick and starts a chat. --speed fast keeps only the ones estimated at 30+ tok/s.