Levelwise
English
C# and .NET basics

Performance and Span

In busy code, each small object on the Heap gives the GC more work. Measure first, then fix only the hot spot. Tools like Span and ArrayPool help you work without copies and without new objects.

Not reviewedWritten with AI helpReading time: 14 minExample of a supplier price file in a shopC# and .NET 10 code

Author: bezzad

The problem: short pauses in the price service

Suppliers send the price and stock of items to our shop. Each message is one line of text: the item id, the price and the stock, separated by commas. Our service parses tens of thousands of lines each second like this:

public static PriceUpdate ParseWithSplit(string line)
{
    var parts = line.Split(',');
    return new PriceUpdate(
        parts[0],
        decimal.Parse(parts[1], CultureInfo.InvariantCulture),
        int.Parse(parts[2], CultureInfo.InvariantCulture));
}

Most of the time the speed is good. But every few seconds, everything stops for a short time. The dotnet-counters tool shows that the number of GCs is very high.

Why are many allocations slow?

Step by step:

  1. The Split method creates one array and three new strings on the Heap for each line.
  2. Tens of thousands of lines per second means hundreds of thousands of objects per second. All of them are garbage a few moments later.
  3. Each new object is created in generation 0. At this speed, generation 0 fills up very quickly.
  4. Each time it fills up, the GC runs. A generation 0 GC is cheap, but when it runs thousands of times, the total time is large.
  5. A few objects that are alive during the GC move to higher generations. Sometimes a full and expensive GC is also needed.

The result: the less garbage we make, the less work the GC does.

line.Split(',')P-907,129.50,12string[3]"P-907""129.50""12"Four new objects in memory, for each lineReadOnlySpan<char>P-907,129.50,12[0..5][6..12][13..]Only three windows on the same memory. No copy and no new object
The Split method creates four new objects for each line. Span only points to pieces of the same original string.

Measure first

The most important rule of Performance is: do not guess. Most of the time, the place we think is slow is not slow.

1. Measuredotnet-counters2. Find hot spotdotnet-trace3. Small changeOnly that spot4. CompareBenchmarkDotNetIf it is not better, undo the change. Simple code has value
Optimization is a cycle. Each change must be proven with a number.
  1. With the dotnet-counters tool, see the overall state: the number of GCs, the Heap size, CPU use.
  2. With the dotnet-trace tool or PerfView, find which method has most of the time or most of the allocations. We call it the hot spot (Hot Path).
  3. Change only that spot. Keep the rest of the code simple.
  4. With the BenchmarkDotNet library, compare before and after.

The BenchmarkDotNet library

This library runs the code many times, does the warm-up, and gives a statistical result. The MemoryDiagnoser attribute also shows the amount of allocation:

using BenchmarkDotNet.Attributes;
using BenchmarkDotNet.Running;

BenchmarkRunner.Run<ParseBenchmarks>();

[MemoryDiagnoser]
public class ParseBenchmarks
{
    private const string Line = "P-907,129.50,12";

    [Benchmark(Baseline = true)]
    public PriceUpdate WithSplit() => PriceParser.ParseWithSplit(Line);

    [Benchmark]
    public PriceUpdate WithSpan() => PriceParser.ParseWithSpan(Line);
}

Always run the program in Release mode:

dotnet run -c Release
Do not measure with Stopwatch. A simple loop with Stopwatch does not take into account JIT warm-up, system noise, and code removed by the compiler. The number it gives is usually misleading.

What is Span?

The Span type is a window on a piece of memory. The memory can be an array, a string, or even memory on the Stack. Span keeps only two things: where it starts and how long it is.

  • The ReadOnlySpan type is read-only. We use this type for strings.
  • Slicing a Span (with a range, like two dots in a row) copies nothing. It only creates a smaller window.
  • Most Parse methods in .NET have a version that takes a ReadOnlySpan.

Now the same parse, without Split:

public readonly record struct PriceUpdate(string Sku, decimal Price, int Stock);

public static PriceUpdate ParseWithSpan(ReadOnlySpan<char> line)
{
    int first = line.IndexOf(',');
    var sku = line[..first];

    var rest = line[(first + 1)..];
    int second = rest.IndexOf(',');

    var price = decimal.Parse(rest[..second], CultureInfo.InvariantCulture);
    var stock = int.Parse(rest[(second + 1)..], CultureInfo.InvariantCulture);

    return new PriceUpdate(sku.ToString(), price, stock);
}

In this version, no array is created, and no strings for the price and the stock are created either. We create only one string for the item id, because we really need it.

Limits of Span

The Span type is a ref struct. This means it can only live on the Stack. The reason is that it may point to Stack memory, and that memory is not valid anymore after the method ends. So:

  1. It cannot be a field of a class.
  2. A lambda cannot capture it.
  3. It cannot cross an await. The variables of an async method are kept on the Heap during an await.
  4. For these cases, use Memory. The Memory type stays on the Heap, and whenever you need one, you get a Span from its Span property.

Reusing a buffer with ArrayPool

Our service also downloads big price files from suppliers. If we create a new 64 KB array for each file, these arrays also become garbage. Arrays of 85 thousand bytes and bigger are even worse: they go directly to the Large Object Heap and are cleaned only by a full GC.

ArrayPool<byte>.SharedArrays ready to borrowRead the price filewith a borrowed arrayRentReturnNo new array is created
Borrow the array, use it and give it back. The pool keeps it for the next job.
public static async Task CopyAsync(Stream input, Stream output, CancellationToken ct)
{
    byte[] buffer = ArrayPool<byte>.Shared.Rent(64 * 1024);
    try
    {
        int read;
        while ((read = await input.ReadAsync(buffer, ct)) > 0)
            await output.WriteAsync(buffer.AsMemory(0, read), ct);
    }
    finally
    {
        ArrayPool<byte>.Shared.Return(buffer);
    }
}
Two important points: The array that Rent gives may be bigger than the size you asked for. So always work with the real number of bytes read. And after Return, never use that array again. Other code may have it now.

stackalloc for small buffers

For a small, temporary buffer, you can take memory on the Stack. Nothing is created on the Heap:

Span<char> buffer = stackalloc char[32];
if (price.TryFormat(buffer, out int written, "N0", CultureInfo.InvariantCulture))
    writer.Write(buffer[..written]);

Only for small, fixed sizes. The Stack is small, and if it fills up, the whole process crashes.

Other common sources of allocation

  • Joining strings in a loop. Each time a new string is created. Use StringBuilder.
  • Creating a List without an initial capacity. When you know the count, give the capacity. Otherwise the List grows and copies its inner array several times.
  • Turning a struct into an object (Boxing). For example, putting an int in a variable of type object or of an interface type.
  • Using LINQ and lambdas in a hot spot. Each step, and each lambda that captures a variable, creates a small object. In normal code this does not matter. In a hot spot, a simple loop is better.
  • An async method that most of the time finishes without waiting. For example, it answers from the cache 99 percent of the time. Here ValueTask removes the allocation of the Task type. But await a ValueTask only once.

Important rules

  1. Measure first. Without numbers, do not optimize.
  2. Optimize only the hot spot. Span code is not more readable than simple code. Keep the rest of the code simple.
  3. Compare before and after with BenchmarkDotNet. In Release mode and with MemoryDiagnoser.
  4. Always give a borrowed buffer back in finally. And after giving it back, do not use it.
  5. Use stackalloc only for small, fixed sizes. Never with a size that comes from the user.
  6. Look at GC settings last. Fix the code first, then the settings.

Common mistakes

Mistake Result The right way
Optimizing without measuring Complex code, with no real benefit. Find the hot spot first.
Measuring with Stopwatch in Debug mode Misleading numbers. The BenchmarkDotNet library in Release mode.
Using Span everywhere in the code Hard-to-read code, with no benefit. Only in the hot spot.
Working with the whole rented array Extra and old data is read. Only up to the real count.
Using the array after Return The data of other code gets broken. Give it back in finally, then never again.
Using stackalloc with a big size The Stack fills up and the process crashes. A small, fixed size, or ArrayPool.

When to use these tools?

Good fit

  • Code that runs thousands of times per second.
  • Measurement has shown that allocation or GC is the problem.
  • Reading and parsing text, working with files and the network, base libraries.

Bad fit

  • A normal endpoint that spends most of its time waiting for the database.
  • Code that runs a few times a day.
  • When you have not measured anything yet.

Summary in six lines

  1. Each new object on the Heap gives the GC more work. In busy code this adds up.
  2. Measure first and find the hot spot. Then change only that.
  3. Compare before and after with BenchmarkDotNet and MemoryDiagnoser in Release mode.
  4. The Span type is a window on memory. Slicing it copies nothing.
  5. The Span type lives only on the Stack. For fields and async code, use Memory.
  6. Borrow big buffers with ArrayPool and give them back in finally.