Follow us
Breaking
Tech Services

Migrating teams to Ollama local LLM runtimes

Teams migrating from OpenAI API can use Ollama for secure local inference. While Ollama delivers 62 tokens per second for single users, vLLM outperforms it at scale, reaching 920 tokens per second with 50 concurrent users.

Share

Local deployment with Ollama

Ollama provides the most flexible path for teams migrating from cloud APIs to local inference. The company, led by Jeff Morgan and Michael Chiang, raised $65 million in Series B funding to help developers run open-weight models on their own hardware. This follows a previous $15 million Series A led by Benchmark’s Peter Fenton. The tool essentially did for AI what Docker and Docker Desktop did for cloud. Ollama serves 8.9 million developers every month and holds 176,000 stars and 17,000 forks on GitHub. Most teams adopt this tool to prevent prompt leaking because they host the platform on VMs under their direct control. This move eliminates monthly subscriptions and the risk of sending proprietary information to cloud providers. Ollama provides a pull-and-run experience for many open-source models, including Meta’s Code Llama. You already know that cloud-based LLMs often require constant internet connectivity and high usage fees. The software works on macOS, Linux, and Windows. Ollama provides access to larger models via its neocloud service.

Performance comparison with vLLM

vLLM serves high-throughput production workloads better than Ollama because of its continuous batching engine. While Ollama delivers 62 tokens per second for a single user with a quantized model, vLLM reaches 71 tokens per second using FP16. At 50 concurrent users, Ollama plateaus at roughly 155 total tokens per second, while vLLM maintains approximately 920 total tokens per second because its engine aggregates multiple requests into unified GPU operations.

Metric Ollama (Llama 3.1 8B) vLLM (Llama 3.1 8B)
Single-User Throughput 62 tok/s 71 tok/s
50-User Throughput 155 tok/s 920 tok/s

Ollama suits teams with fewer than 20 users who prioritize flexibility over massive scale. In contrast, vLLM handles large-scale deployments where many users access the model simultaneously. A benchmark using an NVIDIA RTX 4090 and an AMD Ryzen 9 7500X showed a massive performance gap between these two tools as concurrency increased. The test used Llama 3.1 8B with Q4_K_M quantization for Ollama and FP16 for vLLM using 64 GB of DDR5 RAM. For the Llama 3.1 8B model, the difference between 62 and 71 tokens per second remains small for one person, but the difference becomes extreme when many users hit the system at once. For DeepSeek-R1-Distill-Llama-8B at 50 concurrent users, Ollama reached 142 tokens per second, while vLLM sustained 840 tokens per second. One test recorded 58 tokens per second for DeepSeek-R1-Distill-Llama-8B during single-user runs. One test showed Ollama had a time-to-first-response of 45ms for Llama 3.1 8B, while vLLM recorded 82ms. Can a team balance these speed differences across a large workforce?

Implementing the local stack

Ollama exposes a REST API on port 11434 to facilitate integration with other tools. Teams running Ollama in Docker achieve reproducible environments for their specific workflows. Developers use Modelfiles to customize model behavior, add system prompts, and adjust parameters. A user pulls a model with the ollama pull command and uses ollama run to start it. The ollama list command manages storage, while the ollama ps command monitors loaded models. The ollama stop command frees RAM immediately by unloading unused models. Users require at least 8 GB of RAM, though 32 GB or more helps when running the largest models. For larger models like DeepSeek v4, users require about 128 GB of RAM on high-end systems. The macOS app runs Ollama as a background service automatically. On Linux, the installer configures a systemd service for Linux users. The software supports building AI applications with LangChain and other developer tools.

Share

Technewsdaily

Senior tech writer covering AI, gadgets and cybersecurity. Breaking down the news that matters, every day.