Messaging for the modern edge: NATS explained
NATS provides high-performance messaging through Core NATS and JetStream, offering 1 to 5 ms latency for persistent messages. This lightweight system serves as a low-memory alternative to Apache Kafka for real-time communication and IoT device fleets.
The NATS messaging hierarchy
NATS provides a messaging system that splits functionality into Core NATS and JetStream. Core NATS operates as a high-performance, fire-and-forget system that delivers messages with at-most-once semantics. If a subscriber is offline when a message arrives, the system drops that message. This design allows Core NATS to achieve tens of millions of messages per second per node with microsecond latency. I see this as a useful tool for real-time monitoring or heartbeats where losing a single data point does not break the system. NATS uses subject-based addressing, meaning clients connect via a URL and subscribe to specific strings rather than using IP addresses or DNS. If you already understand the basics of pub/sub, you will recognize how NATS handles location transparency. This allows you to connect to NATS and then talk to anything else that is also connected to the system without knowing its specific location.
JetStream adds a persistence layer on top of the core pub/sub model. It uses the Raft consensus algorithm to handle replication and metadata. JetStream supports at-least-once and exactly-once delivery through consumer acknowledgments and deduplication based on message IDs. You can use it for streaming, work queues, or key-value stores. It also supports message replay, which allows you to re-read data from a specific point in time. JetStream provides different replay policies, such as "instant" which delivers messages without delay, or "original" which delivers them at the speed they were captured. You can also administratively pause or unpause consumer message delivery to coordinate rolling updates of your client applications.
NATS supports many communication patterns beyond simple publish and subscribe. You can implement fan-in, fan-out, and request-reply models. The system also supports scatter-gather patterns through its M to N communications. If you require strict ordering, you can implement a partitioned consumer approach using hash-based partitioning to ensure messages with the same key always go to the same partition. The NATS 2.0 release introduced superclusters, which function as clusters of clusters across different regions, and it uses round-trip delay time to find the lowest latency NATS cluster in the supercluster when routing all clients.
Performance and delivery trade-offs
Comparing NATS JetStream to Apache Kafka or RabbitMQ reveals differences in throughput and operational overhead. Kafka relies on a distributed commit log and uses ZooKeeper to handle clustering and failover. This makes Kafka a complex system to operate. NATS avoids this by using a single binary and the Raft protocol for internal coordination. Liftbridge can also provide Kafka-like log API semantics to NATS, acting as a "voicemail" to the "dial tone" of NATS Core. Liftbridge uses an immutable commit log, which allows multiple consumers to read from the same log and enables message replay for event sourcing. This approach lets you evaluate multiple logging providers simultaneously by providing a way to tee your data.
| Feature | NATS JetStream | Apache Kafka | RabbitMQ |
|---|---|---|---|
| Throughput (msgs/sec) | 200,000 – 400,000 | 500,000 – 1,000,000+ | 50,000 – 100,000 |
| Latency (persistence) | 1 – 5 ms | 10 – 50 ms | 5 – 20 ms |
| Delivery Guarantee | Exactly once | Exactly once | At least once |
| Min. RAM (Production) | 4 GB | 16 GB | 8 GB |
Kafka achieves high throughput through batching, but this batching increases latency to between 10 and 50 milliseconds. NATS JetStream maintains lower latency between 1 and 5 milliseconds when using persistence. If you need to scale a fleet of low-powered IoT devices, the lightweight nature of NATS provides a clear advantage. You should consider if your existing infrastructure can handle the high memory requirements of a Kafka cluster, especially when your application requires the sub-millisecond latency that NATS provides for in-memory operations. NATS JetStream implements R3 clustering with RAFT consensus for metadata and message replication across nodes. You can configure stream replicas on a per-stream basis to meet your specific durability needs.
Market limits and migration costs
I find the argument about NATS’s ceiling compelling. The developers focused on a niche market of lightweight, real-time communication, which limits its mainstream adoption. Because NATS uses its own protocol, you cannot use Kafka Connectors or existing AMQP tools without rewriting your client logic. This creates a high migration cost that many enterprises refuse to pay. I would skip NATS if your project requires deep integration with the massive Kafka ecosystem or existing data pipelines.
The NATS server is written in Go, and the garbage collector can become a bottleneck when the system reaches extreme loads. This performance limitation suggests that NATS might struggle to compete with systems written in Rust or C++ when pushing the absolute limits of hardware. Some users also reported abnormal memory growth in JetStream during specific scenarios, which makes capacity planning difficult.
NATS 2.0 introduced accounts and decentralized security using NKEYS and TLS 1.3. These features provide logical isolation, allowing you to separate different tenants within a single system. You can define how data flows between accounts using services and streams. Users control access via specific permissions to ensure accounts only reach the subjects and data they need. The 2.14 release allows for downsampling messages via scheduled republishing. You should look at the user base if you worry about stability. Companies like Alibaba, Baidu, Walmart, and Ericsson use NATS in production. Financial services like Mastercard and Capital One also rely on it.
Will the growing ecosystem of NATS-specific tools ever close the gap with Kafka?