CI/CD and Safe Deployment
Every code change is built and tested automatically and makes one fixed package. The same package goes to the environments step by step. With safe deploy methods and step-by-step database changes, deploy becomes a daily task with no stress.
Author: bezzad
The problem: the Friday night deploy
Our shop team deploys once a month. Each time, one person builds the app on their own laptop, copies the files to the server and runs the database migration by hand. The result:
- There are too many changes. A month of work is in one deploy. If something breaks, finding the cause is hard.
- Each time is a bit different. A step is forgotten, or another SDK version is installed on the laptop.
- Everyone is afraid. So they delay the deploy, and the changes grow even more.
The solution is to make deploy small, automatic and repeatable.
The idea: an automatic production line (Pipeline)
Each time someone pushes code, an automatic Pipeline runs.
- The CI part (Continuous Integration). The code is built and all tests run. If something is broken, we know in a few minutes, not a month later.
- The CD part (Continuous Delivery). After testing, a package (Image) is built and goes to the environments step by step.
We have two similar names that are different:
- Continuous Delivery. Every version is ready to go to Production. But a person presses the final button.
- Continuous Deployment. There is no button. Every change that passes all tests goes to Production automatically.
For the second step we need very good tests and strong monitoring. Most teams start with the first step.
Build once, use everywhere
An important rule: build the package only once. The same package goes to the test environment, then to the main environment.
Why?
- If you build again for each environment, a package may get a new version in between.
- So what reaches Production is not exactly what you tested.
- With one fixed package and a commit tag, you know exactly what runs. Going back to the previous version also means only running the previous Image.
So settings (database address, secrets) are not inside the package. Each environment gives its own settings from outside.
Safe deploy methods
When the new version is ready, how do we put it in place of the old version so the user feels nothing?
- The Rolling method. Servers or Pods are replaced one by one. It is simple and does not need many extra servers. But for a few minutes two versions run together. This is the default method in Kubernetes.
- The Blue-Green method. The new version starts fully next to the old version. Then traffic moves to the new version at once. Going back is very fast, because the old version is still running. But for a while you need double the resources.
- The Canary method. First a small percent of users see the new version. If errors and response time are good, the percent goes up slowly. If there is a problem, only a few people have seen it.
The database: the hardest part of deploy
In the Rolling and Canary methods, the old and new versions of the code work with one database at the same time. So at every moment the database must be compatible with both versions.
Example: the customers table has a “full name” column. We want to turn it into two columns, “first name” and “last name”. If one migration deletes the old column right at the start:
- Some Pods still have the old code and read the old column.
- The database gives a “column does not exist” error, and the user sees a 500 error.
- If the new version has a bug, going back is also not possible. The old code needs a column that no longer exists.
The solution is the Expand and Contract pattern. We do the big change in several small deploys:
- Expand. Add the new columns. Allow them to be empty, so the old code does not break.
- Write to both. The new code fills both the old column and the new columns. Fill old records in small batches, not with one big command that locks the table for a long time.
- Read from the new. When all data is filled, the code reads only from the new columns.
- Contract. A few days later, when you are sure you will not go back to the old version, delete the old column.
Code
The Pipeline file with GitHub Actions
name: orders-api
on:
push:
branches: [main]
pull_request:
jobs:
build-test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-dotnet@v4
with:
dotnet-version: '10.0.x'
- run: dotnet restore
- run: dotnet build --no-restore -c Release
- run: dotnet test --no-build -c Release
publish-image:
needs: build-test # only after all tests pass
if: github.ref == 'refs/heads/main'
runs-on: ubuntu-latest
env:
IMAGE: registry.example.com/shop/orders-api:${{ github.sha }}
steps:
- uses: actions/checkout@v4
- uses: docker/login-action@v3
with:
registry: registry.example.com
username: ${{ secrets.REGISTRY_USER }}
password: ${{ secrets.REGISTRY_PASSWORD }}
- run: docker build -t "$IMAGE" .
- run: docker push "$IMAGE" # tagged with the commit, built once
A few points:
- For each Pull Request only build and test run. Building the Image is only for the main branch.
- Secrets come from the safe secrets section, not from the file itself.
- The Image tag is the commit id. So you always know which code is in which environment.
The first Expand step in EF Core
public partial class AddFirstAndLastName : Migration
{
protected override void Up(MigrationBuilder migrationBuilder)
{
migrationBuilder.AddColumn<string>(
name: "FirstName", table: "Customers", nullable: true);
migrationBuilder.AddColumn<string>(
name: "LastName", table: "Customers", nullable: true);
// No DropColumn("FullName") here.
// It is removed in a later release, after all code stops using it.
}
protected override void Down(MigrationBuilder migrationBuilder)
{
migrationBuilder.DropColumn(name: "FirstName", table: "Customers");
migrationBuilder.DropColumn(name: "LastName", table: "Customers");
}
}
Running the migration apart from the app
If each Pod runs the migration when it starts, several Pods may do it at the same time. It is better for the migration to be a separate step in the Pipeline. The EF Core tool can build all migrations into one executable file (a bundle):
# In the pipeline: build the bundle once
dotnet ef migrations bundle --self-contained -r linux-x64 -o efbundle
# Before the new code is deployed: run it once
./efbundle --connection "$ORDERS_DB_CONNECTION"
Important rules
- Many small changes. The smaller each deploy, the lower the risk and the easier it is to find a problem.
- A fast first step. Build and unit tests should take a few minutes. If the Pipeline takes an hour, programmers work around it.
- A failed step means stop. Do not ignore a red test. Fix a flaky test, do not run it again until it turns green.
- Build once. The same package for all environments.
- Secrets only in secrets. No secret in the code, the Pipeline file or the Image.
- Every migration must be compatible with the previous code. Deleting or renaming a column happens only in the last step of Expand and Contract.
- Look after deploy. Check the error rate and response time after every deploy. If they get worse, go back quickly.
Common mistakes
| Mistake | Result | The right way |
|---|---|---|
| Building again for each environment | What was tested is different from what was released. | Build once and tag with the commit. |
| Deleting or renaming a column in one migration | A 500 error in the middle of deploy, and no way back. | Expand and Contract over several deploys. |
| Running the migration when each Pod starts | Several Pods run the migration at the same time. | A separate step in the Pipeline, for example with a bundle. |
| A secret in the Pipeline file | Anyone with access to the code repository sees the secret. | The secrets section of the CI tool. |
| Running a flaky test again until it turns green | Nobody trusts the tests anymore. | Find the cause and fix the test. |
| No rollback plan | When there is a problem, the team does not know what to do. | Go back to the previous Image in one step. |
Which deploy method?
Simple and enough
- The Rolling method is enough for most services, as long as the database change is compatible with both versions.
- The Canary method is for sensitive services like payment, where errors are expensive.
Be careful
- The Blue-Green method for a system that does not have enough resources for two full versions.
- Continuous Deployment without strong tests and monitoring. Errors go straight to the customer.
Summary in six lines
- In CI every change is built and tested automatically. In CD the same result goes to the environments step by step.
- Build the package once and tag it with the commit id. Settings come from outside.
- Rolling is simple, Blue-Green has a fast rollback, Canary makes the risk small.
- With a Feature Flag, deploy is separate from release.
- The database must be compatible with both the old and the new code. Use Expand and Contract.
- Small changes, a fast Pipeline, secrets in secrets, and a look at monitoring after every deploy.