What Is a Feature Flag? Turning Features On and Off Without a Redeploy

What is a feature flag?
A feature flag (also called a feature toggle) is a conditional switch that lets you turn a software feature on or off without redeploying code. The application checks a configuration value to decide whether to run the feature. As a result, the feature can sit in production switched off, then go live with a single setting change.
Think about the wiring in a house. The cables already run inside the walls, and the lamps hang in place. However, the light stays dark until someone presses the switch. A feature flag is that switch: the code is already in place, and you choose later when to turn the light on.
In this guide we answer the question what is a feature flag at a conceptual level. We also look at the main types, the risks, and the cleanup habit that keeps flags from turning into clutter.
What is a feature flag in terms of deploy versus release?
In a traditional workflow, two events happen at the same moment. For example, you push code to the server, and users see the new feature immediately. So a feature flag pulls these two events apart.
Deploying means moving code to the production environment. Releasing means exposing a feature to real users. With a flag, you deploy first while the feature stays off, so users notice nothing. Then you make the release decision later, when you are ready.
This separation has three practical effects:
- Deployments become routine, because they carry no visible change for users.
- Also, product teams pick the release moment independently of the engineering calendar.
- If something breaks, you turn the flag off instead of rolling back code.
So the old habit of avoiding Friday deployments loses some of its force. Instead, the risky moment is no longer the code landing on the server. It is the moment the feature turns on, and that decision now belongs to you. If container-based delivery is new to you, our guide to Docker and containers gives useful background.
How does a feature flag work in code?
The logic is simple. First, the code asks a flag for a value, the flag returns one, and the code follows one of two paths. While the new path is off, the old behavior keeps running exactly as before.
A few building blocks sit behind this. The flag key names the flag. A default value is the safe fallback for the moment the real value fails to load. Next, the evaluation context holds the information needed to decide, such as a user ID, a country or a plan type. Finally, rules define which context receives which value.
The flow looks like this:
- First, the code asks for the flag by key and supplies a default value.
- Then the evaluator receives the user or request context.
- Next, it checks the rules you defined.
- A value comes back: on, off, a string or a number.
- The code runs the new or the old path based on that value.
The rules can live in a config file, a database or a dedicated management panel. However, what matters is that the value can change without redeploying the code. For example, you tick a box in the panel, and within seconds your servers see the new value.
One design habit also helps. Make the flag decision in a single place, not scattered across the code. Ask the question once, then pass the answer to the relevant component as behavior. Therefore, removing the flag later is a small and predictable change.
What types of feature flags exist?
Not every flag does the same job. However, the classification in Martin Fowler's feature toggle article groups flags into four categories by purpose and lifespan. Many teams use these four as shared vocabulary.
- Release toggle: Hides unfinished work in production code. It is usually short lived.
- Experiment toggle: Splits users into cohorts to run an A/B test. You remove it once the result is clear.
- Ops toggle: Changes system behavior at runtime. Some of them stay as permanent kill switches.
- Permissioning toggle: Opens a feature to a specific group, such as beta users or a higher plan. It can live for years.
The table below puts the four side by side. Also, the lifespan and decision columns help you plan when each flag should disappear.
| Type | Main purpose | Typical lifespan | How the decision changes |
|---|---|---|---|
| Release toggle | Hide unfinished features | Short | Mostly fixed per deployment |
| Experiment toggle | Run A/B tests | Short to medium | Per request, per user |
| Ops toggle | Disable a problem component | Short, or long for a kill switch | Changes quickly at runtime |
| Permissioning toggle | Give a group special access | Long | Per request, per user |
This split matters in practice. Each type needs a different owner, a different cleanup schedule and a different testing approach.
Are feature flags only on or off, or can they hold other values?
No, flags are not always binary. The most common form is a boolean: on or off. However, many systems also support string, number and object values.
This flexibility, however, serves different needs. For example, a string value can tell a screen which layout variant to show. A number can carry an adjustable threshold, such as how many products appear per page. An object can deliver several settings in one bundle.
Also, multi-value flags are especially useful in experiments. You can manage three or four variants with a single flag instead of opening one flag per variant.
That said, simple is valuable. A plain on or off flag is the easiest to understand, test and delete. So choose a multi-value flag only when you truly need variants. The OpenFeature documentation also defines four typed evaluation methods and expects a default value on every call.
When do release and permissioning flags help?
A release flag lets a team merge into the main codebase often. Instead of keeping a large feature on a separate branch for weeks, you merge small pieces into the main line. The pieces stay hidden behind the flag, so unfinished work never reaches users.
Teams call this trunk based development. In other words, everyone works on one main line. Merge conflicts shrink, because branches never drift far apart.
A permissioning flag solves a different problem: who sees what. For example, you can open a new reporting screen to your internal team first, then to a beta list, and finally to a higher plan. Besides, you do not need a separate code version for each group.
Keep one point in mind. Permissioning flags live long, so they turn into product rules. Label them separately from temporary flags and document them.
How do experiment flags and kill switches differ?
An experiment flag assigns the same user to the same cohort every time. For example, half of your visitors see the new cart layout and the other half see the old one. You then compare the behavior of the two groups. So do not decide before the result is statistically meaningful. Our A/B test calculator can help with that check.
A kill switch works like an emergency breaker. When an external service slows down, you switch off the secondary feature that depends on it. The main flow keeps working, and only the extra feature goes dark for a while. People call this graceful degradation.
The two differ in speed. For instance, you plan an experiment flag in advance and watch it patiently. In contrast, you must flip a kill switch within seconds during a crisis. Therefore, access to the kill switch controls should be clear, and the switch should keep working when the system is under load.
What are gradual rollouts and canary releases?
A gradual rollout opens a feature to a small percentage first, then grows step by step instead of exposing everyone at once. First comes your internal team, then a small slice of users, then everybody. At each step you watch the metrics and move on only if nothing looks wrong.
The canary name comes from miners who once carried canaries to test the air. Because the bird showed danger before people felt it, the image fits. Likewise, a small user slice is the canary for a new feature: if something fails, only a few people feel it.
Let us clear up a common mix-up. A canary deployment works at the infrastructure level and sends part of the traffic to the new version. A feature flag works inside the application and opens a feature per user. You can use both together.
Also, monitoring is essential for gradual rollouts. If you cannot see error rates, response times and business metrics, raising the percentage is guesswork. We cover the monitoring side in our article on observability and OpenTelemetry.
One more detail: split the percentage by a stable user ID. Otherwise the same person may see a feature appear and vanish after a page refresh. Stable bucketing therefore protects the user experience and keeps your measurements clean.
What are the benefits of using feature flags?
The value of a feature flag comes from cutting risk into small pieces. Instead of one big release, you make small decisions that you can each undo on their own.
- Smaller risk: If something breaks, you turn the flag off and skip a new deployment.
- Early feedback: You try the feature with a few real users first.
- Easier merging: Unfinished work stays hidden in the main code, so branches stay short.
- Calendar freedom: You pick a campaign or event day independently of engineering.
- Controlled experiments: Decisions rest on data, not guesses.
In addition, a flag gives teams a shared language. A product manager says "open it to ten percent," and an engineer applies it with one setting. The short cycles we describe in our article on agile project management with Scrum and Kanban become safer with flags.
Still, these benefits do not arrive automatically. If you do not manage flags regularly, the load builds up instead of the value. Therefore, treat a flag as a tool that needs maintenance, not as a one time feature.
What risks do feature flags bring?
Every flag adds a new branch to your code. For that reason, a flag that is easy to open can become an expensive burden over time. These are the risks we see most often:
- Flag debt: Flags that finished their job but stayed in the code make it harder to read.
- Test complexity: As the flag count grows, the number of combinations to try grows fast.
- Access and security: If everyone can flip flags, someone may switch off a live feature by mistake.
- Latency: Asking a remote service on every request can slow down responses.
- Wrong context: If you send an incomplete user ID, your experiment groups blur together.
None of these risks cancels the idea of a flag. However, each one needs a planned countermeasure. If you handle user data, also read our guide to building a GDPR compliant website.
How do you manage flag permissions and security?
A flag panel is effectively a remote control for your live system. Therefore, treat access to it as seriously as access to the production database. Because of that, one wrong click can switch off a working feature for everyone.
A simple but effective permission approach includes the following:
- Limit who can change production flags to a small group.
- Record who changed what, when and to which value.
- Keep development, test and production in separate flag sets.
- Ask for a second person's approval on critical flags.
Also keep the evaluation context small. For most decisions, a user ID or a plan type is enough, and you do not need names, emails or phone numbers. As a result, the flag system does not turn into an unnecessary store of personal data.
Finally, plan a backup path for emergency flags such as kill switches. If the main flag service goes down, the thing you want to switch off may be the flag service itself.
How do you prevent flag debt and clean up old flags?
Fowler's article suggests treating flags like inventory: every flag has a carrying cost. Cleanup, then, is not a luxury for the end of the project. Instead, it is a promise you make when you create the flag.
A simple cleanup policy works like this. First you name an owner, then you set a target end date:
- Write an owner and a target removal date when you create the flag.
- Label the flag type, so its expected lifespan is clear.
- Add the removal task to the backlog together with the flag.
- When the feature is fully live, delete the old path and the flag together.
- Review the flag inventory at regular intervals.
- Add an automatic warning or a failing test, a so called time bomb, for flags past their date.
Some teams also cap the number of active flags. To add a new one, you must delete an old one. The rule sounds harsh, but it really stops flag debt.
Keep permanent flags, such as permissioning toggles and kill switches, outside this deletion cycle. Label them as long lived and track them in a separate list.
Who decides about a flag, and how do roles split?
Technically, an engineer adds the flag. However, the decision to turn it on often belongs to someone else. So define the roles up front. Otherwise the flag becomes a pile of settings that nobody owns.
In practice, a split like this works:
- Engineer: Adds the flag to the code, tests it, sets the default value and takes responsibility for deleting it.
- Product manager: Decides which user slice sees it and when, and defines the success measure.
- Quality owner: Confirms the tests for both the on and the off state.
- Operations or support: Knows the runbook for emergency flags such as kill switches.
In small teams, one person may hold several of these roles. Even so, write the roles down, because nobody should argue about "who can turn it off?" during an incident.
In addition, make flag work visible in sprint planning. Adding a flag and deleting it are both tasks, and both take time.
How do feature flags compare with branching strategies and blue green deployment?
All three approaches try to manage risk, but they work at different layers. Therefore, they complement each other instead of competing. The table below shows the difference.
| Approach | What it separates | How you roll back | Main risk | Best fit |
|---|---|---|---|---|
| Feature flag | Deployment from release | Turn the flag off | Flag debt, test combinations | Gradual opening, experiments |
| Long lived feature branch | Code from the main line | Do not merge the branch | Conflicts at late merge | Short, small changes |
| Blue green deployment | Old and new environment | Switch traffic back to the old one | Cost of running two environments | Switching a whole release at once |
A branching strategy decides where code lives. Blue green deployment builds two identical environments and moves traffic from one to the other. A feature flag, in contrast, targets one feature inside a single environment. For a refresher on branches, see our Git and GitHub guide.
Mature teams usually combine all three. For instance, they switch the release with blue green and keep the risky feature off with a flag.
When should you not use a feature flag?
Adding a flag to every change is also a mistake. A small text fix or a style tweak does not need one, because the flag adds more complexity than it saves.
Be careful in these cases:
- Database schema changes: A flag turns code off, but it does not reverse data. Plan a backward compatible, step by step migration.
- Security fixes: Therefore, do not make a change that closes a hole optional.
- Simple fixed settings: For values that never change, environment configuration is enough.
- Very short lived work: For a change that ships and finishes the same day, a flag is just extra load.
The general rule is this: if the flag has no clear question to answer, skip it. If the question is "do we want to open this gradually, test it or turn it off fast?", a flag makes sense.
What is OpenFeature and why does it matter for feature flags?
OpenFeature is an open, vendor neutral standard API for feature flagging. According to its official documentation, it is a project under the Cloud Native Computing Foundation (CNCF). You can check its current maturity level on the CNCF project page.
The goal is simple: you should not have to rewrite application code when you change your flag tool. The standard has these building blocks:
- Evaluation API: The interface your application uses to ask for a flag.
- Evaluation context: The carrier of the data needed for the decision.
- Provider: The layer that translates the standard call into the flag system you chose.
- Hooks: Add behavior such as validation, logging or telemetry to the evaluation lifecycle.
- Events: Signal changes in the provider state or the flag configuration.
The specification also sets a safe design rule. Typed evaluation (boolean, string, number, object) always takes a default value. If evaluation fails, the application does not crash and the default value comes back. For sources, the OpenFeature introduction and the CNCF project page are good starting points.
Should you build flag management yourself or use a ready made service?
At its simplest, flag logic starts with a config file. For a small team, an environment variable or a JSON file can be enough. Still, this approach breaks down as needs grow.
These signs suggest it is time for a more mature management layer:
- Non technical teammates want to flip a flag without changing code.
- You need percentage based rollouts or per user rules.
- Also, you must keep a record of who changed what and when.
- Several applications and languages share the same flags.
Broadly, there are three options: a simple tool you write yourself, an open source management system, or a paid hosted service. Each has a different maintenance load, cost and flexibility. For that reason, we do not recommend a specific product. Instead, we give selection criteria.
Your criteria should include latency impact, offline behavior, access and audit logs, data location and standards support. Moreover, if you use a standard layer such as OpenFeature, changing this choice later gets much easier. For pricing and feature comparisons, check each provider's current official documentation.
What is a feature flag in a real scenario?
To make the idea concrete, here is a sample scenario that answers what is a feature flag in daily work. An e-commerce team is building a new cart page. Over several weeks, the code merges into the main line in small pieces, yet customers keep seeing the old page because the cart flag is off.
On launch day, the team follows this order:
- They turn the flag on for the internal team and test the cart by hand.
- Then they open it to a small customer slice and watch errors and abandonment.
- If nothing looks wrong, they widen the slice step by step.
- If a problem appears, they turn the flag off and return to the old cart.
- Once everybody uses the new cart, they delete the old code and the flag.
Nobody redeploys anything along the way. Moreover, if the sales team wants the cart to wait until a campaign day, that is still just a setting.
In a flow like this, rollbacks must be safe. To make sure repeating an action causes no harm, also read our article on idempotency and idempotent APIs.
How do you test code that sits behind a feature flag?
After the question what is a feature flag, testing is the next topic. Adding a flag enlarges the test space. In effect, each flag means the code can behave in two ways. Therefore, you must answer the question "which states are we testing?" explicitly.
Fowler's advice is simple: at least test the intended production state and the fallback state. Most flags are independent, so you do not need to try every combination. If two flags interact, test them together.
- Run the main scenario with the flag on and off.
- Verify the default behavior when the flag value fails to load.
- Check that the same user lands in the same group every time.
- Finally, confirm that the old path is really gone after you remove the flag.
Also, tying your pre release checks to a list reduces mistakes. Our website pre launch testing and QA checklist gives you a ready skeleton for this.
How do feature flags work with AI assisted code and observability?
Generating code quickly with AI means code you have not verified yet gets closer to production. Putting such changes behind a flag gives you a safe buffer. We described our view in our article on vibe coding, so here we only cover the role of the flag.
Observability, on the other hand, is the flag's eyes. If error rate, latency or business metrics change when you open a flag, you need to see it. When you attach the flag value to logs and traces as a label, you can find which flag state started the problem.
So OpenFeature hooks exist for exactly this job. They let you add logging or telemetry at evaluation time. In other words, the link between the flag decision and the measurement stays intact.
In short, the trio works like this: the flag limits risk, monitoring shows the impact, and code review protects the quality bar.
What is a practical feature flag checklist?
First, answer these questions before taking a flag live. If you cannot say yes to all of them, the flag is not ready yet.
- Does the flag have an owner and a clear name?
- Is the default value on the safe side?
- Did you label its type and expected lifespan?
- Did you test both the off and the on state?
- Is it clear who can change the flag, and is there a record?
- Do you know which metrics to watch once it is on?
- Is the rollback step written down?
- Did you add the deletion task to the backlog?
- Can the decision work without personal data in the context?
Move this list to your team wiki and use it for every new flag. That way, flag culture depends on a process, not on one person.
How can our team help you set up feature flags?
At Talha Aslan and team, we recommend making the flag approach part of the architecture from the start of a software project. We settle early which flags are temporary and which are permanent, who manages them and how we monitor them.
Adding flags to an existing application is also possible. Still, starting small with one risky feature is the healthiest path. For details, see our custom software development service.
In short, the answer to what is a feature flag is a release switch that shrinks risk and leaves the decision to you. One last reminder: this article offers general information. Tools, prices and features change, so always check current values in the provider's official documentation.



