Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
47 changes: 47 additions & 0 deletions content/blog/apache-iggy-or-apache-kafka.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
---
title: "Apache Iggy or Apache Kafka: which problems each one is actually for"
description: Where Apache Iggy fits against Apache Kafka, when to reach for each one, and what to expect if you already run Kafka.
author: Justin Mclean
tags: ["kafka", "comparison", "architecture"]
date: 2026-09-23
---

The question I get asked most about Iggy is whether it replaces Kafka. For some things it does and for plenty of others it doesn't, so the useful part is knowing which is which.

## What each one is built for

Kafka is built for scale, and for everything that has grown up around it. The broker runs on the JVM and stores each topic as a partitioned append-only log, replicated to a set of in-sync replicas. KRaft replaced ZooKeeper in 4.0, so a quorum of brokers holds metadata rather than a separate service. Share groups went production-ready in 4.2, so it provides queue semantics as well as log semantics. But what you're really choosing is the ecosystem: Connect, Streams, Schema Registry, every CDC tool and every observability vendor already speak Kafka.

Iggy handles the same streaming work, with lower tail latency and less to run. It's written in Rust and runs thread-per-core on io_uring, with nothing shared between cores, so each partition is owned by a single CPU-pinned shard. It speaks its own binary protocol over TCP, QUIC and WebSocket, with a REST API alongside. Running it means one binary and one config file, with no JVM and nothing else to install. There are SDKs for Rust, Python, Go, Java, C#, Node, C++ and PHP, and it has consumer groups and stores consumer offsets on the server, so consumers work much as they do in Kafka. It's younger than Kafka by a decade, and the current release is 0.9.0.

## Reach for Iggy when

**Operational weight is your main cost.** You're running a single node or a small cluster, and you don't want to operate a JVM cluster: tuning heap and GC, planning partition counts up front, running a rebalancer. Iggy is one binary and one TOML file, with the CLI, Prometheus metrics and OpenTelemetry traces built in. A cluster is the same binary, with every node loading the same TOML file. If you want the Web UI, that's a separate app with its own Docker image.

**Tail latency matters more than ecosystem breadth.** There's no garbage collector, so no GC pauses, and thread-per-core with CPU pinning keeps the tail tight. You'll see the difference in your slowest requests, not in messages per second. Iggy publishes its benchmarks at [benchmarks.iggy.apache.org](https://benchmarks.iggy.apache.org) and ships iggy-bench, so you can run the same tests on your own hardware rather than taking anyone's number.

**The deployment target is constrained.** You're deploying to the edge, to a customer's own hardware, or to a single box where a JVM cluster is the wrong shape. Thread-per-core runs on io_uring, so it needs a recent Linux kernel: 5.19 or newer.

**The network is unfriendly.** QUIC and WebSocket are first-class transports here, with TLS on all three. Kafka has nothing equivalent, and that matters when producers sit on links that drop out or behind proxies that only pass HTTP.

**You're starting fresh and Iggy has the connectors you need.** There are 15 sinks and 4 sources, covering Postgres, S3, Iceberg, ClickHouse, Elasticsearch and InfluxDB among others.

## Reach for Kafka when

**You need a long production record.** Kafka's replication has a decade of production behind it and every failure mode is written up somewhere. Iggy's clustering shipped in 0.9.0. It's based on Viewstamped Replication and tested with deterministic simulation, but that still isn't the same as years of real incidents.

**Your architecture depends on the ecosystem.** You need three connectors and a stream processor, or a schema registry, or an integration a vendor already ships.

**You're running at a scale nobody has run Iggy at.** Someone has to be first, but unless you're willing to do that proving yourself, Kafka is the known quantity.

## If you're already on Kafka

Iggy has a Kafka wire protocol gateway, but it's a work in progress. The plan is to bridge existing Kafka producers and consumers so you can move across gradually without much change on the client side. Until it ships, moving means changing clients.

The realistic pattern is not migration but placement: Kafka where the ecosystem and the scale are, Iggy for the latency-sensitive or footprint-constrained piece that Kafka is awkward for.

## What would change this

This changes as more people run Iggy's clustering in production, when there are more connectors, and when the Kafka gateway is finished.

Until then, Iggy is a good fit for a real and growing set of problems, and Kafka is still the answer for most of the rest. Anyone who tells you it's a straight swap either hasn't run Kafka or hasn't looked at what Iggy ships today.
25 changes: 25 additions & 0 deletions content/docs/introduction/coming-from-kafka.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
---
title: Coming from Apache Kafka
description: "How Apache Kafka's topics, partitions, consumer groups and offsets map to Apache Iggy, and what works differently."
---

If you already know Apache Kafka, most of Apache Iggy will feel familiar. This page maps the Kafka terms you know to Iggy's and lists what works differently. For when to choose one or the other, see [Apache Iggy or Apache Kafka](/blogs/2026/09/23/apache-iggy-or-apache-kafka).

## How the terms map

| Kafka | Iggy | What to know |
|---|---|---|
| No equivalent | Stream | A container for related topics. For example, a `dev` stream and a `production` stream can each have their own `orders` topic. |
| Topic | Topic | Belongs to one stream. |
| Partition | Partition | Where messages are stored, as in Kafka. |
| Consumer group | Consumer group | Each partition goes to one member of the group, and the server rebalances when members join or leave. |
| Offset | Offset | Starts at 0 in each partition, but can skip numbers after a recovery. |
| Retention | Retention policy | Set per topic. Whole segments are deleted once all their messages have expired. |

[Concepts](/docs/introduction/concepts) explains each of these in more detail.

## What works differently

**Iggy has its own protocol.** A Kafka client can't talk to Iggy yet, so you use an [Iggy SDK](/docs/sdk/introduction), the HTTP API or the [binary protocol](/docs/binary-protocol). A Kafka protocol gateway is being built (see [issue #3421](https://github.com/apache/iggy/issues/3421)), but it can't serve consumers yet.

**There is an extra level above topics.** Every topic lives in a stream, so you create a stream before you create its topics.
44 changes: 44 additions & 0 deletions content/docs/introduction/glossary.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
---
title: Glossary
description: "Short definitions of the terms used across the Apache Iggy docs, with links to where each one is explained."
---

A one-line definition of each term, with a link to the page that explains it.

**Append-only log.** An ordered list of messages that only grows at the end. Messages can't be changed once written, and you can read from any earlier point to replay them. See [Concepts](/docs/introduction/concepts#append-only-log).

**Auto-commit.** An option on a poll. The server saves the consumer's position as it returns the messages, before the consumer has processed them. See [Polling messages](/docs/introduction/concepts#polling-messages).

**Connection string.** A single value such as `iggy://user:password@localhost:8090` that tells a client how to connect. The Rust SDK and the SDKs that wrap it (Python, C++ and PHP) accept one. See [Connection strings](/docs/sdk/connection-strings).

**Consumer.** A client that reads messages on its own rather than as part of a consumer group. Its position is kept separately from every other consumer's.

**Consumer group.** A set of consumers that share a topic's partitions, with each partition read by one member at a time. See [Consumer groups](/docs/introduction/concepts#consumer-groups).

**Durability.** A topic setting, `replicated` or `persisted`, that decides what must be done before a write is reported as successful. See [Durability](/docs/server/durability).

**Offset.** The position of a message within a partition, starting at 0. See [Concepts](/docs/introduction/concepts#append-only-log).

**Partition.** Where a topic's messages are actually stored. A topic's messages are spread across its partitions. See [Partition](/docs/introduction/concepts#partition).

**Partitioning.** How a producer picks the partition for a message: a fixed partition, round robin, or a hash of a message key. See [Getting started](/docs/introduction/getting-started).

**Payload.** The body of a message. The server stores it as bytes without interpreting it, so the format is up to you. See [Serialising messages](/docs/introduction/serialising-messages).

**Personal access token.** A token you create on the server and use in place of a username and password, for example in a scraper or a script. See [Connection strings](/docs/sdk/connection-strings).

**Primary and replica.** In a cluster, the primary orders writes and the replicas copy them. If the primary fails, the replicas elect a new one. See [Viewstamped Replication](/docs/clustering/vsr).

**Retention policy.** Topic settings that delete old messages by age or by size. See [Topic options](/docs/server/topic-options).

**Segment.** Part of a partition, stored on disk as a data file and an index file. Retention deletes old segments whole. See [Segment](/docs/introduction/concepts#segment).

**Shard.** One thread in the server that owns a set of partitions. See [Architecture](/docs/introduction/architecture).

**Stream.** A container for related topics. For example, a `dev` stream and a `production` stream can each have their own `orders` topic. See [Stream](/docs/introduction/concepts#stream).

**Topic.** A named set of partitions for one kind of message, such as orders or user events. See [Topic](/docs/introduction/concepts#topic).

**User headers.** Optional key and value pairs sent with a message, kept separate from the payload. See [Serialising messages](/docs/introduction/serialising-messages).

**VSR.** Viewstamped Replication, the protocol Apache Iggy uses to keep replicas in step. See [Viewstamped Replication](/docs/clustering/vsr).
2 changes: 1 addition & 1 deletion content/docs/introduction/meta.json
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
{
"title": "Introduction",
"pages": ["about", "concepts", "architecture", "quickstart", "getting-started"]
"pages": ["about", "concepts", "coming-from-kafka", "glossary", "serialising-messages", "architecture", "quickstart", "getting-started"]
}
68 changes: 68 additions & 0 deletions content/docs/introduction/serialising-messages.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
---
title: Serialising messages
description: "Choose a format for your message payloads, and send more than one kind of message to the same topic."
---

The server stores each message payload as bytes without interpreting it, so you choose the format: plain text, JSON, Protobuf or anything else. This page shows how to send more than one kind of message to the same topic, so the consumer knows how to decode each one.

The examples are in Rust and come from the [examples](https://github.com/apache/iggy/tree/master/examples/rust) in the Apache Iggy repository.

## One kind of message

If every message on a topic has the same shape, send the encoded value as the payload and decode it the same way on the other side. [Getting started](/docs/introduction/getting-started) does this with plain strings.

## A type field in the payload

Wrap each message in an envelope that names its type, and send the envelope as JSON:

```json
{ "message_type": "order_confirmed", "payload": "{\"order_id\":42,\"price\":99.5}" }
```

The consumer decodes the envelope, reads `message_type`, then decodes the inner payload as that type. This works with any client, but every message is decoded twice, and the inner payload is a JSON string inside the envelope. See the [message-envelope example](https://github.com/apache/iggy/tree/master/examples/rust/src/message-envelope).

## A type field in a header

Messages can carry user headers, which are key and value pairs kept separate from the payload. Put the type in a header and send the payload as it is:

```rust
let mut headers = BTreeMap::new();
headers.insert(
HeaderKey::try_from("message_type").unwrap(),
HeaderValue::try_from(message_type).unwrap(),
);

let message = IggyMessage::builder()
.payload(Bytes::from(json))
.user_headers(headers)
.build()
.unwrap();
```

The consumer reads the header and decodes the payload once:

```rust
let payload = std::str::from_utf8(&message.payload)?;
let header_key = HeaderKey::try_from("message_type").unwrap();
let headers_map = message.user_headers_map()?.unwrap();
let message_type = headers_map.get(&header_key).unwrap().as_str()?;

match message_type {
ORDER_CREATED_TYPE => {
let order_created = serde_json::from_str::<OrderCreated>(payload)?;
info!("{:#?}", order_created);
}
// One arm for each message type.
_ => {
warn!("Received unknown message type: {}", message_type);
}
}
```

This is the better choice for Protobuf or any other binary format, because the payload doesn't have to fit inside JSON. See the [message-type example](https://github.com/apache/iggy/tree/master/examples/rust/src/message-headers/message-type).

Header values don't have to be strings. They can also be numbers, booleans or raw bytes, so you can add things like a trace ID next to the type, as the [typed-headers example](https://github.com/apache/iggy/tree/master/examples/rust/src/message-headers/typed-headers) shows. The [message-compression example](https://github.com/apache/iggy/tree/master/examples/rust/src/message-headers/message-compression) uses a header in the same way to record how the payload was compressed.

## Size limits

A payload can be up to 64 MB, and a message's user headers up to 100 KB in total. The whole send request also has to fit the server's request limit, which by default is 64 MiB over TCP and 2 MB over HTTP.
Loading