Trust is earned, not given

A different perspective

2025-07-15 · Projects

Messaging in .NET, part 1: Kafka, or how two programs talk without knowing each other

Part 1from the Messaging in .NET series · 3 parts in all

Sooner or later you write two programs that need to talk. One takes an order from a customer; another one has to charge a card and update inventory. The obvious move is to have the first program call the second one over HTTP. It works, right up until the second program is being deployed, or is slow, or is down. Now the first program has to know whether to retry, how long to wait, and what to do about the order it could not hand over — and you have accidentally made two systems into one system.

This tutorial is for someone who has never used a message broker. We will take that same pair of programs and put a broker between them: a piece of infrastructure whose only job is to hold messages until someone is ready for them. The sample is a small .NET 10 project called KafkaDotNet, and by the end of this article you will have run it, published an order, and watched a background worker process it.

What a broker actually buys you

Think about a post office rather than a phone call. With a phone call, both people must be present at the same moment; if the other end is not answering, there is nothing to say. With mail, the sender drops a letter in a box and gets on with their day. The recipient collects it whenever they like. If they are away for a week, the letter is still there.

That is the whole idea, and it gives you four things for free:

You pay for that with a new moving part to run, and with a new kind of bug: the message that gets processed twice. We will come back to that, because it is the single most important thing to understand about this style of programming.

The one idea behind Kafka: it is a log, not a queue

Most people arrive thinking a broker is a queue: you put a message in, a worker takes it out, it is gone. Kafka is not that. Kafka is an append-only log — closer to a ledger than a queue. Messages are added to the end and they stay there, in order, for as long as the log's retention policy says (typically days). Readers do not remove anything; they just remember how far they have read.

That single design choice explains almost everything else about Kafka. Because consuming does not delete, ten different programs can read the same messages without interfering with each other. Because messages stay, you can fix a bug on Tuesday and re-process last week's orders. And because it is a log, "where am I up to?" is just a number per reader — Kafka calls it an offset.

Three words to learn first

TermPlain meaning
TopicA named log. orders.placed is where orders go. Think of it as a folder in the filing cabinet.
PartitionA topic is split into several logs so that work can happen in parallel. Each partition is still strictly ordered on its own.
Consumer groupA name shared by workers that should split the work between them. Two workers in the same group get different messages; workers in different groups each get everything.

Ordering is per partition, and the partition is chosen by the message's key. That is the one piece of design you have to get right, and the sample makes it explicit:

// Key the order by customer, so every order for one customer lands on the
// same partition and is processed in the order it was placed.
Key = PartitionKeys.ForCustomer(order.CustomerId)

If you key by a random value instead, you get maximum parallelism and no ordering at all. There is no right answer in general — only a right answer for what you are building.

Run the sample

You need Docker and the .NET 10 SDK. Everything below is typed into a terminal; you do not need to read any code first.

git clone https://github.com/bobhuang1/KafkaDotNet
cd KafkaDotNet

# 1. Start a Kafka broker (and a web UI to peek at it)
docker compose up -d

# 2. Start the worker that processes orders - leave this running
dotnet run --project src/KafkaDotNet.OrderProcessor

# 3. In a second terminal: start the API that accepts orders
dotnet run --project src/KafkaDotNet.OrderApi

The first run creates the topics automatically, so there is no setup step. Now send an order:

curl -X POST http://localhost:5080/orders \
  -H 'Content-Type: application/json' \
  -d '{ "customerId": "customer-1",
        "items": [ { "sku": "SKU-001", "quantity": 2, "unitPrice": 19.99 } ] }'

The API answers with an orderId, a partition number and an offset. The interesting output is in the worker's terminal, where you should see something like:

Published OrderPlaced ... for customer-1 to orders.placed[1] at offset 0
Handled order ... for customer-1 on delivery 1
Produced OrderConfirmed to orders.confirmed[2] offset 0 (attempt 0)
Committed orders.placed[1] offset 0 for order ...: Confirmed

Read those four lines as a story: it arrived, it was processed, the result was published, and only then was the offset moved forward. That last line is the part that matters most.

The two pieces of code worth understanding

Writing a message

Producing to Kafka is one call. The important part is not the call — it is the key and the delivery guarantee:

var record = new Message<string, OrderPlaced>
{
    Key = PartitionKeys.ForCustomer(order.CustomerId),   // decides the partition
    Value = order,
};

// Wait for the broker to confirm the write before we tell the customer "yes".
var receipt = await _producer.ProduceAsync(OrderTopics.Placed, record, cancellationToken);

The sample's producer is configured with Acks.All (the broker only confirms once every replica has the message) and EnableIdempotence (the producer's own internal retries cannot write the same message twice). Those two settings are the difference between "we think it arrived" and "we know it arrived".

Reading messages, and the one line everyone gets wrong

var record = consumer.Consume(TimeSpan.FromMilliseconds(500));
if (record is null) continue;          // nothing new; read again

// ... do the actual work here ...

consumer.Commit(record);               // <-- only AFTER the work succeeded

An offset is a bookmark, and Commit moves the bookmark. So the order of those two lines is the entire delivery guarantee:

The catch has a name, and it is the sentence to remember from this whole article: at-least-once means your work can happen twice, so the code doing it must be safe to run twice (charge the card once, by checking whether that order was already charged, rather than by assuming you are only ever called once).

What happens when things go wrong

Two kinds of failure, and the sample treats them differently on purpose.

A business "no". The SKU is out of stock. Retrying will not help, ever. So the worker publishes an OrderRejected to a different topic and moves on. Nothing loops.

A temporary "not now". The payment provider timed out. This might work in five seconds. The worker pushes the order onto a retry topic with a not-before timestamp on it, and a second consumer holds it back until that time passes. If it keeps failing, after a few attempts the order is parked on a dead-letter topic — a place for messages that a human needs to look at, rather than messages that quietly disappeared.

One thing worth noticing: Kafka has no built-in delay. There is no "deliver this in five seconds" button. The sample builds it out of a timestamp and a whole topic, which is more work than you would expect. Remember that when we get to RabbitMQ in part three.

Try the interesting cases

The sample ships a demo endpoint that walks through every branch at once:

curl -X POST 'http://localhost:5080/orders/demo?count=6'

Order zero carries a SKU that is always out of stock (it gets rejected), order one is deliberately over the large-order threshold (it fails once and then succeeds on the retry), and the rest succeed immediately. Watch the worker's log and you will see all four outcomes in about ten seconds, which is much better than reading about them.

Beginner mistakes to avoid

Where to go next

You now know the shape of every message-based system: something appends, something reads, something decides where the bookmark goes, and something happens when the work fails. What changes between brokers is how much of that they hand you.

In part 2 we build the identical pipeline on Redis, which many teams already run for caching — and see which of Kafka's guarantees you get for free and which ones you have to build yourself. In part 3 we implement the pipeline a third time on RabbitMQ and compare the two head to head.