The Mathematics Of Software Development Efficiency
Introduction
In software development, the developers have been discussing the advantages and disadvantages of different development strategies, particularly trunk-based development versus feature-based development. However, all of those are plain words and opinions; there is no formal mathematical analysis that quantifies the efficiency of these approaches.
In this blog post, I will provide a formal mathematical analysis to quantify the efficiency of trunk-based development and feature-based development, and demonstrate that trunk-based development is generally more efficient in modern complicated software development scenarios.
Software Development Efficiency
With some assumptions and simplifications, I will derive the mathematical formulas to quantify the merge conflict resolution effort and test cost associated with both trunk-based and feature-based development strategies.
Feature-Based Development
Feature-based development is a software development approach where new features are developed in isolated branches and integrated into the main codebase only after they are complete. This approach allows developers to work independently on different features without affecting the stability of the main branch.
Conceptually, the first branch created from the main branch serves as the main branch for the feature project. Multiple developers can then create additional feature branches from it to work dependently from each other on their individual tasks. This feature project’s main branch can itself be developed using either trunk-based or feature-based strategies. Thus, whenever development branches are created, developers can choose between trunk-based and feature-based development strategies. In the worst scenario, it can become a very deeply nested tree of branches, making integration and coordination exponentially complex.
To simplify our analysis, we assume that no further branches are created from the feature branch branched off from the main branch and all the developers working on the feature are so cooperative that the patches are being stacked sequentially and smoothly without conflicts.
Suppose there are $N$ patches, $P_{1}, P_{2}, \ldots, P_{N}$, being stacked, and $P_{i}$ can have dependencies on the previous patches $P_{1}, P_{2}, \ldots, P_{i-1}$. If the feature branch ever gets rebased, in the worst case, every single patch will have merge conflicts that need to be resolved. The size of each patch is denoted by $S_{1}, S_{2}, \ldots, S_{N}$. In the meanwhile, the master branch also evolves. Whenever there is a conflict between the master branch and the feature branch, the feature branch will rebase onto the master branch and resolve the conflicts accordingly.
Without loss of generality, we assume that resolving a merge conflict for a patch of size $S_{i}$ requires effort proportional to $S_{i}$ and bounded by $\mathcal{O}(S_{i})$. We denote that the number of times each patch $P_{i}$ ecounters merge conflicts during feature development by $C_{i}$.
The total effort required to resolve all merge conflicts during feature development can be expressed as:
$$
\begin{align}
E &= C_{1} \mathcal{O} \left( \sum_{i=1}^{N} S_{i} \right) + C_{2} \mathcal{O} \left( \sum_{i=2}^{N} S_{i} \right) + \ldots + C_{N} \mathcal{O} \left( S_{N} \right)
\end{align}
$$
Because $\mathcal{O} \left( \sum_{i=j}^{N} S_{i} \right) \geq \mathcal{O} \left( S_{j} \right)$ for all $j$, we have the following lower bound for the total effort:
$$
\begin{align}
E &\geq C_{1} \mathcal{O} \left( S_{1} \right) + C_{2} \mathcal{O} \left( S_{2} \right) + \ldots + C_{N} \mathcal{O} \left( S_{N} \right) \\
&= \sum_{i=1}^{N} C_{i} \mathcal{O} \left( S_{i} \right)
\end{align}
$$
The feature branch will sometimes also have to rebase because of requiring additional updates or fixes to its dependencies from the main branch. Even if there is no direct conflict, there are test cost for each patch whenever there is a rebase.
Without loss of generality, we assume that the test cost for a patch of size $S_{i}$ is proportional to $S_{i}$ and bounded by $\mathcal{O}(S_{i})$. We denote that the number of times each patch $P_{i}$ requires rebasing because of dependency updates from the main branch by $R_{i}$.
The total test cost associated with rebasing due to dependency updates can be expressed as:
$$
\begin{align}
T &= R_{1} \mathcal{O} \left( \sum_{i=1}^{N} S_{i} \right) + R_{2} \mathcal{O} \left( \sum_{i=1}^{N} S_{i} \right) + \ldots + R_{N} \mathcal{O} \left( \sum_{i=1}^{N} S_{i} \right) \\
&= \left( \sum_{i=1}^{N} R_{i} \right) \mathcal{O} \left( \sum_{i=1}^{N} S_{i} \right) \\
&= \sum_{i=1}^{N} \left( R_{i} \mathcal{O} \left( \sum_{i=1}^{N} S_{i} \right) \right) \\
&\geq \sum_{i=1}^{N} R_{i} \mathcal{O} \left( S_{i} \right)
\end{align}
$$
Note that the test cost associated with each rebase is always $\mathcal{O} \left( \sum_{i=1}^{N} S_{i} \right)$ because the rebase can have an impact on all patches in the feature branch.
If there are nested branches, the rebase and testing costs can become exponentially more expensive, which we will skip the analysis here. This is definitely something we should avoid in practice.
Trunk-Based Development
Trunk-based development, on the other hand, encourages developers to integrate their changes into the main branch frequently, often multiple times a day. This approach minimizes integration issues and promotes continuous collaboration among team members. It also allows for faster feedback and more efficient use of development resources, as the main branch always reflects the current state of the project.
Multiple developers will create branches from the main branch to work independently from each other on their individual tasks. There might be some additional effort to consolidate all these changes, which can be large or small depending on how the software is architected and how the changes were made. Because previously in feature-based development we assumed all the developers working on the feature are so cooperative that the patches are being stacked sequentially and smoothly without conflicts, in trunk-based development the consolidation effort is zero.
Suppose there are $N$ patches $P_{1}^{\prime}, P_{2}^{\prime}, \ldots, P_{N}^{\prime}$, each patch is an independent branch created from the main branch. If the branch ever gets rebased, only one patch can have the possibility of having merging conflicts. The size of each patch is denoted by $S_{1}^{\prime}, S_{2}^{\prime}, \ldots, S_{N}^{\prime}$. In the meanwhile, the master branch also evolves. Whenever there is a conflict between the master branch and a patch, the branch associated with that patch will rebase onto the master branch and resolve the conflicts accordingly.
Without loss of generality, we assume that resolving a merge conflict for a patch of size $S_{i}^{\prime}$ requires effort proportional to $S_{i}^{\prime}$ and bound by $\mathcal{O} \left( S_{i}^{\prime} \right)$. We denote that the number of times each patch $P_{i}^{\prime}$ encounters merge conflicts during feature development by $C_{i}^{\prime}$.
The total effort required to resolve all merge conflicts during feature development can be expressed as:
$$
\begin{align}
E^{\prime} &= C_{1}^{\prime} \mathcal{O} \left( S_{1}^{\prime} \right) + C_{2}^{\prime} \mathcal{O} \left( S_{2}^{\prime} \right) + \ldots + C_{N}^{\prime} \mathcal{O} \left( S_{N}^{\prime} \right) \\
&= \sum_{i=1}^{N} C_{i}^{\prime} \mathcal{O} \left( S_{i}^{\prime} \right)
\end{align}
$$
Each branch will sometimes also have to rebase onto the main branch to incorporate some of the latest changes required to its dependencies.
Without loss of generality, we assume that the test cost for a patch of size $S_{i}^{\prime}$ is proportional to $S_{i}^{\prime}$ and bounded by $\mathcal{O}(S_{i}^{\prime})$. We denote that the number of times each patch $P_{i}^{\prime}$ requires rebasing because of dependency updates from the main branch by $R_{i}^{\prime}$.
The total test cost associated with rebasing due to dependency updates can be expressed as:
$$
\begin{align}
T^{\prime} &= R_{1}^{\prime} \mathcal{O} \left( S_{1}^{\prime} \right) + R_{2}^{\prime} \mathcal{O} \left( S_{2}^{\prime} \right) + \ldots + R_{N}^{\prime} \mathcal{O} \left( S_{N}^{\prime} \right) \\
&= \sum_{i=1}^{N} R_{i}^{\prime} \mathcal{O} \left( S_{i}^{\prime} \right)
\end{align}
$$
Trunk-Based Development VS Feature-Based Development
With the mathematical analysis above, directly comparing the effort and cost of trunk-based development versus feature-based development remains still infeasible, because the patches used in the two approaches, $P_{i}$ and $P_{i}^{\prime}$, are different.
However, if we assume $C_{i} = C_{i}^{\prime}$, $S_{i} = S_{i}^{\prime}$, and $R_{i} = R_{i}^{\prime}$, for all $i$, which is sometimes not too bad in practice, we can find the following relationships:
$$
\begin{align}
E &\geq E^{\prime} \\
T &\geq T^{\prime}
\end{align}
$$
This suggests that, under the assumptions made, trunk-based development is at least as efficient as feature-based development in terms of both conflict resolution effort and testing cost.
Minimizing Trunk-Based Development Consolidation Effort
One of the assumptions made in the previous analysis is that the consolidation effort and cost is zero, meaning that when all the patches are independently merged into the main branch, the project is complete. This is theoretically impossible, unless all the features from all different patches are truly independent from each other, which is rarely the case in a collaborative development project. But this additional consolidation effort and cost can be minimized through careful planning, modular design, and effective dependency management.
If the software architecture and contribution process involves modular design, clear dependency management, and well-defined interfaces, for each branch in the trunk-based development, the development will create or modify only their assigned modules or components, without conflicting with changes made by other developers. The new features implemented in the branches should be verified and be set to disabled when merged into the main branch. As a consequence, when each branch is merged, it should have no effect to the existing functionality of the main branch. After all the branches have been merged, it then comes to the consolidation phase, where the merged features are being enabled and tested in combination to ensure the overall system works correctly and efficiently. In a well-structured project, the work involved in this consolidation phase should only be toggling a few feature flags followed by system testing.
Conclusions
In a complex software development project where multiple developers and extremely complicated dependencies are involved, assuming a good software architecture, such as modular design and clear dependency management, trunk-based development is definitely superior to feature-based development in terms of minimizing development effort and cost. To some extent, trunk-based development is the best practice, as it reflects careful problem decomposition and planning, leading to more efficient and manageable development processes.
References
The Mathematics Of Software Development Efficiency
https://leimao.github.io/blog/The-Mathematics-Of-Software-Development-Efficiency/