Messaging in .NET, part 1: Kafka, or how two programs talk without knowing each other
Part 1from the Messaging in .NET series · 3 parts in all
Sooner or later you write two programs that need to talk. One takes an order from a customer; another one has to charge a card and update inventory. The obvious move is to have the first program call the second one over HTTP. It works, right up until the second program is being deployed, or is slow, or is down. Now the first program has to know whether to retry, how long to wait, and what to do about the order it could not hand over — and you have accidentally made two systems into one system.
This tutorial is for someone who has never used a message broker. We will take that same pair of programs and put a broker between them: a piece of infrastructure whose only job is to hold messages until someone is ready for them. The sample is a small .NET 10 project called KafkaDotNet, and by the end of this article you will have run it, published an order, and watched a background worker process it.
What a broker actually buys you
Think about a post office rather than a phone call. With a phone call, both people must be present at the same moment; if the other end is not answering, there is nothing to say. With mail, the sender drops a letter in a box and gets on with their day. The recipient collects it whenever they like. If they are away for a week, the letter is still there.
That is the whole idea, and it gives you four things for free:
- Nobody has to be available at the same time. The sender writes; the receiver reads. Never simultaneously.
- A crash does not lose the message. It is already sitting in the broker, not in the memory of a process that just died.
- Retrying is the broker's problem, not yours. You stop writing code that loops on a failing HTTP call.
- You can add a third reader. A notifications service can read the same orders without you changing the sender at all.
You pay for that with a new moving part to run, and with a new kind of bug: the message that gets processed twice. We will come back to that, because it is the single most important thing to understand about this style of programming.
The one idea behind Kafka: it is a log, not a queue
Most people arrive thinking a broker is a queue: you put a message in, a worker takes it out, it is gone. Kafka is not that. Kafka is an append-only log — closer to a ledger than a queue. Messages are added to the end and they stay there, in order, for as long as the log's retention policy says (typically days). Readers do not remove anything; they just remember how far they have read.
That single design choice explains almost everything else about Kafka. Because consuming does not delete, ten different programs can read the same messages without interfering with each other. Because messages stay, you can fix a bug on Tuesday and re-process last week's orders. And because it is a log, "where am I up to?" is just a number per reader — Kafka calls it an offset.
Three words to learn first
| Term | Plain meaning |
|---|---|
| Topic | A named log. orders.placed is where
orders go. Think of it as a folder in the filing cabinet. |
| Partition | A topic is split into several logs so that work can happen in parallel. Each partition is still strictly ordered on its own. |
| Consumer group | A name shared by workers that should split the work between them. Two workers in the same group get different messages; workers in different groups each get everything. |
Ordering is per partition, and the partition is chosen by the message's key. That is the one piece of design you have to get right, and the sample makes it explicit:
// Key the order by customer, so every order for one customer lands on the
// same partition and is processed in the order it was placed.
Key = PartitionKeys.ForCustomer(order.CustomerId)
If you key by a random value instead, you get maximum parallelism and no ordering at all. There is no right answer in general — only a right answer for what you are building.
Run the sample
You need Docker and the .NET 10 SDK. Everything below is typed into a terminal; you do not need to read any code first.
git clone https://github.com/bobhuang1/KafkaDotNet
cd KafkaDotNet
# 1. Start a Kafka broker (and a web UI to peek at it)
docker compose up -d
# 2. Start the worker that processes orders - leave this running
dotnet run --project src/KafkaDotNet.OrderProcessor
# 3. In a second terminal: start the API that accepts orders
dotnet run --project src/KafkaDotNet.OrderApi
The first run creates the topics automatically, so there is no setup step. Now send an order:
curl -X POST http://localhost:5080/orders \
-H 'Content-Type: application/json' \
-d '{ "customerId": "customer-1",
"items": [ { "sku": "SKU-001", "quantity": 2, "unitPrice": 19.99 } ] }'
The API answers with an orderId, a partition number and an offset. The
interesting output is in the worker's terminal, where you should see something
like:
Published OrderPlaced ... for customer-1 to orders.placed[1] at offset 0
Handled order ... for customer-1 on delivery 1
Produced OrderConfirmed to orders.confirmed[2] offset 0 (attempt 0)
Committed orders.placed[1] offset 0 for order ...: Confirmed
Read those four lines as a story: it arrived, it was processed, the result was published, and only then was the offset moved forward. That last line is the part that matters most.
The two pieces of code worth understanding
Writing a message
Producing to Kafka is one call. The important part is not the call — it is the key and the delivery guarantee:
var record = new Message<string, OrderPlaced>
{
Key = PartitionKeys.ForCustomer(order.CustomerId), // decides the partition
Value = order,
};
// Wait for the broker to confirm the write before we tell the customer "yes".
var receipt = await _producer.ProduceAsync(OrderTopics.Placed, record, cancellationToken);
The sample's producer is configured with Acks.All (the broker only confirms
once every replica has the message) and EnableIdempotence (the producer's own
internal retries cannot write the same message twice). Those two settings are the
difference between "we think it arrived" and "we know it arrived".
Reading messages, and the one line everyone gets wrong
var record = consumer.Consume(TimeSpan.FromMilliseconds(500));
if (record is null) continue; // nothing new; read again
// ... do the actual work here ...
consumer.Commit(record); // <-- only AFTER the work succeeded
An offset is a bookmark, and Commit moves the bookmark. So the order of
those two lines is the entire delivery guarantee:
- Commit before the work: if you crash in the middle, the bookmark says "done" and the message is skipped forever. The order vanishes. This is at-most-once.
- Commit after the work: if you crash in the middle, the bookmark never moved, so the message is delivered again and processed again. This is at-least-once — and it is what you almost always want.
The catch has a name, and it is the sentence to remember from this whole article: at-least-once means your work can happen twice, so the code doing it must be safe to run twice (charge the card once, by checking whether that order was already charged, rather than by assuming you are only ever called once).
What happens when things go wrong
Two kinds of failure, and the sample treats them differently on purpose.
A business "no". The SKU is out of stock. Retrying will not help, ever.
So the worker publishes an OrderRejected to a different topic and moves on.
Nothing loops.
A temporary "not now". The payment provider timed out. This might work
in five seconds. The worker pushes the order onto a retry topic with a
not-before timestamp on it, and a second consumer holds it back until that
time passes. If it keeps failing, after a few attempts the order is parked on a
dead-letter topic — a place for messages that a human needs to look at,
rather than messages that quietly disappeared.
One thing worth noticing: Kafka has no built-in delay. There is no "deliver this in five seconds" button. The sample builds it out of a timestamp and a whole topic, which is more work than you would expect. Remember that when we get to RabbitMQ in part three.
Try the interesting cases
The sample ships a demo endpoint that walks through every branch at once:
curl -X POST 'http://localhost:5080/orders/demo?count=6'
Order zero carries a SKU that is always out of stock (it gets rejected), order one is deliberately over the large-order threshold (it fails once and then succeeds on the retry), and the rest succeed immediately. Watch the worker's log and you will see all four outcomes in about ten seconds, which is much better than reading about them.
Beginner mistakes to avoid
- Treating the broker as a database. Kafka keeps messages for a while, not forever. If you need the data permanently, write it somewhere that says so.
- Assuming "exactly once". End-to-end exactly-once delivery is not something you switch on. Design the work to be safe when repeated.
- Not setting a key when order matters. Without one, messages for the same customer can land on different partitions and overtake each other.
- Committing on a timer. Auto-commit is convenient and quietly switches you to at-most-once. Turn it off and commit where you mean it.
- More workers than partitions. Consumers in a group cannot outnumber partitions; the extra ones sit idle.
Where to go next
You now know the shape of every message-based system: something appends, something reads, something decides where the bookmark goes, and something happens when the work fails. What changes between brokers is how much of that they hand you.
In part 2 we build the identical pipeline on Redis, which many teams already run for caching — and see which of Kafka's guarantees you get for free and which ones you have to build yourself. In part 3 we implement the pipeline a third time on RabbitMQ and compare the two head to head.