llmpm
by llmpm
A package manager for open-source LLMs - install, run and serve models from the command line
llmpm is an MIT-licensed command-line package manager for open-source AI models, styled after pip and npm. It installs and runs text, vision, image, speech-to-text and text-to-speech models pulled from HuggingFace Hub, Ollama and Mistral AI, and can serve any of them behind an OpenAI-compatible HTTP API.
llmpm is a command-line package manager for open-source large language models, published under the MIT licence and maintained in the llmpm/llmpm-dev repository on GitHub. It deliberately borrows the pip and npm mental model: 'llmpm install <repo>' downloads a model, 'llmpm run <repo>' opens an interactive session with automatic model-type detection, and 'llmpm list', 'llmpm info' and 'llmpm uninstall' manage what is on disk. 'llmpm search' queries the HuggingFace Hub, 'llmpm benchmark' runs evaluations, and 'llmpm push' uploads a model back to HuggingFace. Installation is offered three ways - 'pip install llmpm' (recommended), 'npm install -g llmpm', and a Homebrew tap - and every route bootstraps an isolated virtual environment at ~/.llmpm/venv so heavy ML backends never touch the system Python. It operates in two modes: global, storing models in ~/.llmpm/models/, and local, where the presence of an llmpm.json file redirects downloads into a project-scoped .llmpm/models/ directory for reproducible per-project setups. Backends are selected by artifact type: GGUF files run through llama.cpp, Transformers checkpoints through the HuggingFace libraries, with diffusion and audio pipelines alongside them, all installable selectively. 'llmpm serve' starts a single HTTP server that can host several models at once, exposing OpenAI-compatible chat completions plus image, speech-to-text and text-to-speech endpoints and a browser chat UI at /chat. The project requires Python 3.9 to 3.13, declares no runtime dependencies in its npm package, and version 3.1.4 was published to PyPI on 12 March 2026. It was introduced on Hacker News as 'Show HN: Llmpm - NPM for LLMs' on 9 March 2026 by author sarthaksaxena, who cited the 'friction we kept seeing when trying to run models locally', and it publishes model rankings at llmpm.co/rankings.
Individual developers and small ML teams who run open-weight models locally and want one CLI across HuggingFace, Ollama and Mistral instead of three separate workflows and three separate Python environments.
One command installs a model and one more serves it behind an OpenAI-compatible endpoint, with every backend isolated from the system Python.
At a Glance
- Category
- Developer Tools
- Pricing
- Free
- Target Market
- Enterprise Developers, Data Scientists, CTOs
- Deployment
- Open-source, Self-hosted
- Founded
- 2026
Key Features
- ✓pip, npm or Homebrew installation
Three install paths that all bootstrap an isolated virtual environment so backends never touch the system Python.
- ✓Multi-registry model install
Pulls models from HuggingFace Hub, Ollama and Mistral AI through one consistent install command.
- ✓OpenAI-compatible serve command
llmpm serve runs one HTTP server hosting several models with chat, image, speech-to-text and text-to-speech endpoints.
- ✓Project-local model scoping
An llmpm.json file redirects downloads into .llmpm/models so each repository pins its own model set.
- ✓Automatic backend selection
GGUF artifacts route to llama.cpp while Transformers checkpoints use HuggingFace libraries, chosen without any user configuration.
- ✓Built-in benchmarking
llmpm benchmark runs evaluations locally and feeds the public model rankings published at llmpm.co/rankings.
- ✓Browser chat UI
The serve command exposes a chat interface at /chat so a model can be tried without extra tooling.
Capabilities
Use Cases
- •Local model evaluation
A developer installs several candidate models and benchmarks them on one machine before committing to a hosting bill.
- •Offline application development
Point an OpenAI-compatible SDK at the local serve endpoint so application code works without network access or API keys.
- •Reproducible project setup
Commit an llmpm.json alongside the code so every contributor pulls the same model versions into the repository.
- •Multimodal prototyping
Run vision, speech-to-text and text-to-speech models from the same CLI while sketching a multimodal feature.
- •Publishing a local fine-tune
llmpm push uploads a locally adapted model back to HuggingFace without setting up a separate publishing workflow.
Ideal For
Best For
- ✓Trying open-weight models locally without hand-managing llama.cpp, Transformers and diffusion environments
- ✓Standing up a local OpenAI-compatible endpoint for application development against open models
- ✓Pinning models per project via an llmpm.json file so a repository's model set is reproducible for every contributor
- ✓Comparing candidate models with the built-in benchmark command and the public rankings
- ✓Working across text, vision, image, speech-to-text and text-to-speech models from one interface
Not Ideal For
- ✗Production enterprise deployment - the project is early and small, with 16 GitHub stars and a single fork at the current snapshot, and no SLA, support contract or security certification
- ✗Teams already standardised on Ollama, vLLM or LM Studio, which cover most of the same local-serving ground with far larger communities and longer track records
- ✗GPU-cluster or multi-tenant serving, where a single-host CLI offers nothing over vLLM, Text Generation Inference or Ray Serve
Deployment
Market Analysis
Pros
- ✓MIT licence, no account requirement and no declared runtime dependencies in the npm package mean very low adoption friction
- ✓One CLI spans three model registries and four modalities, replacing several single-purpose tools
- ✓Project-local llmpm.json scoping makes a repository's model set reproducible, which most local runners do not do
- ✓The isolated ~/.llmpm/venv keeps heavy ML backends out of the system Python, avoiding a common local-inference failure mode
Cons
- ✗Very small adoption: 16 GitHub stars and 1 fork, and the Show HN launch on 9 March 2026 drew only 6 points and 2 comments
- ✗No independent review coverage exists - G2, Capterra, TrustRadius and Product Hunt returned nothing for this product
- ✗It reads as a single-maintainer project with no published governance, roadmap, release cadence or support commitment
- ✗The vendor site llmpm.co returned HTTP 402 to unattended requests during research, so its published claims could not be verified directly
- ✗Heavy overlap with Ollama and LM Studio, both far more widely deployed, leaves the switching case resting on multi-registry support and project scoping alone
Pricing
Open source (MIT)
$0
- ✓Full CLI under the MIT licence
- ✓Unlimited local models
- ✓OpenAI-compatible serve endpoint and browser chat UI
- ✓No account or API key required
llmpm is free and MIT-licensed with no paid tier and no account requirement, and its npm package declares no runtime dependencies. The real cost is local hardware - GPU or RAM sufficient for the models you pull - plus whatever disk the model cache consumes under ~/.llmpm/models, or under a project's .llmpm/models directory when an llmpm.json is present.
Security & Compliance
Connect
Sources
This page was written from 5 sources, 5 on domains other than llmpm.co.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Raindrop
AI agent monitoring that catches silent production failures: Sentry for AI agents
GitLab Duo Agent Platform
Agentic AI across the whole GitLab DevSecOps lifecycle: planning, coding, code review, CI/CD and security agents under one governance model
CodeRabbit
AI code review and agentic change management for teams shipping human- and machine-written code
Qodo
Agentic code review and governance layer for teams shipping AI-generated code at scale