Original title: "a16z: Why is chain performance difficult to measure?"
Original author: Joseph Bonneau, a16z
Original translation: Katie Gu, Odaily Planet Daily
Performance and scalability are much-discussed challenges in the crypto space, related to L1 projects and L2 solutions. However, we do not have standardized metrics or benchmarks. Data is often reported inconsistently and incompletely, which makes it difficult to accurately compare projects and often obscures what is most important in practice.
We need a more nuanced and thorough approach to measuring and comparing performance - breaking performance into multiple components and comparing trade-offs on multiple data axes. This article defines the basic meaning of blockchain performance and outlines its challenges, and provides guidelines and key principles to keep in mind when evaluating blockchain performance.

First, scalability and performance have standard computer science meanings that are often misused in blockchain. Performance measures what a system is currently able to achieve. As we will discuss below, performance metrics may include transactions per second or median transaction confirmation time. Scalability, on the other hand, measures the ability of a system to improve performance by adding resources.
The distinction is important: if defined correctly, many ways to improve performance do not improve scalability at all. A simple example is using a more efficient digital signature scheme, such as BLS signatures, which are about half the size of Schnorr or ECDSA signatures. If Bitcoin switched from ECDSA to BLS, the number of transactions per block might increase by 20-30%, improving performance overnight. But we can only do this once - there's no more space-efficient signature scheme to switch to (BLS signatures can also be aggregated to save more space, but that's another one-time trick).
It's possible to use other one-time tricks in blockchains (like SegWit), but you need a scalable architecture to achieve continuous performance improvements, where adding more resources improves performance over time. This is also the traditional way of thinking about many other computer systems, such as building a network server. With a few common tricks, you can build a server that runs fast. But you ultimately need a multi-server architecture where you keep adding additional servers to keep up with growing demand.
Understanding this distinction also helps avoid common category errors found in statements like “Blockchain X is highly scalable, it can process Y transactions per second.” The second statement may be impressive, but it is a performance metric, not a scalability metric, and does not refer to the ability to increase performance by adding resources.
Scalability inherently requires the exploitation of parallelism. In the blockchain space, L1 scaling appears to require forks or something similar to forks. The basic concept of forks, which split the state into blocks so that different validators can process them independently, matches the definition of scalability very well. There are more options at L2 that allow for the addition of parallel processing, including off-chain channels, Rollup servers, and sidechains.
Typically, blockchain system performance is evaluated along two dimensions, latency and throughput. Latency measures how quickly individual transactions are confirmed, while throughput measures the total rate of transactions over time. These axes apply to both L1 and L2 systems, as well as many other types of computer systems (such as database query engines and web servers).
But both latency and throughput are complicated to measure and compare. Furthermore, individual users don’t really care about throughput (which is a system-wide metric). What they really care about is latency and transaction fees. More specifically, that their transactions get confirmed as quickly and cheaply as possible. While many other computer systems are also evaluated on a cost/performance basis, transaction fees are a new performance axis for blockchain systems that doesn’t really exist in traditional computer systems.
Latency seems simple at first: how long does it take for a transaction to get confirmed? But there are several different ways to answer this question.
First, we can measure latency between different points in time, which can give different results. For example, do we start measuring latency when the user hits the “Submit” button locally, or when the transaction hits the mempool? Do we stop counting time when a transaction is in a proposed block, or when a block is confirmed by one or six subsequent blocks?
The most common approach is to measure the time from when a user first broadcasts a transaction to when the transaction is reasonably “confirmed” from the perspective of a validator. Of course, different merchants may have different acceptance criteria, and even a single merchant may have different criteria based on the amount of the transaction.
The validator-centric approach ignores a few important things in practice. First, it ignores latency on the peer-to-peer network (how long does it take from a client broadcasting a transaction until most nodes hear it?) and client latency (how long does it take to prepare a transaction on the client’s local machine?). Client latency is likely to be very small and predictable for simple transactions like signing Ethereum payments, but can be very important for more complex cases like proving that a shielded Zcash transaction is correct.
Even if we standardize the time window over which we attempt to measure latency, the answer is that it depends. To date, no cryptocurrency system has ever provided fixed transaction latency. A basic rule of thumb is this:Latency is a distribution, not a number.
The networking research community has long understood this. Special emphasis is placed on the “long tail” of the distribution, since even 0.1% of transactions (or web server queries) with high latency can severely impact end users.
In blockchains, confirmation latency can vary for many reasons:
Batching: Most systems batch transactions in some way, e.g., into blocks on most L1 systems. This will result in variable latency, as some transactions will have to wait until the batch fills up. Others may get lucky and join in last. These transactions are confirmed immediately and do not experience any additional latency.
Variable congestion: Most systems experience congestion, meaning more transactions are posted than the system can immediately process. The degree of congestion can vary when transactions are broadcast at unpredictable times, or when the rate of new transactions varies over the day or week, or in response to external events like a hot NFT launch.
Consensus layer differences: Confirming transactions at L1 typically requires a distributed set of nodes to reach consensus on a block, which can add variable latency regardless of congestion. Proof-of-Work systems find blocks at unpredictable times. PoS systems can also add various delays (for example, if not enough nodes are online to form a committee in a round, or if a change of opinion is needed in response to a leader crash).
For these reasons, a good guideline should be:
Statements about latency should show the distribution of confirmation times, rather than a single number like the mean or median.
While summary statistics like the mean, median, or percentiles provide part of the picture, accurately evaluating a system requires considering the entire distribution. In some applications, if the latency distribution is relatively simple, average latency can provide good insight. But in cryptocurrencies, this is almost never the case. Typically, confirmation times are very long.
Payment channel networks (like the Lightning Network) are a good example of this. This is a classic L2 scaling solution, and these networks provide very fast payment confirmations most of the time, but occasionally they require channel resets, which can increase latency by orders of magnitude.
Even if we have good statistics about the exact latency distributions, they may change as systems and system requirements change. Also, it is not always clear how latency distributions compare between competing systems. For example, suppose a system confirms transactions with latency uniformly distributed between 1 and 2 minutes (mean and median of 90 seconds). If a competing system accurately confirms 95% of transactions in 1 minute, and the other 5% in 11 minutes (mean 90 seconds, median 60 seconds), which system is better? The answer may be that some applications prefer the former, and some prefer the latter.Finally, it is important to note that in most systems not all transactions are given the same priority. Users can pay more to get higher priority, so in addition to all of this, latency also depends on the transaction fee paid. In summary:
Latency is complex. The more data reported, the better. Ideally, the full latency distribution should be measured under varying congestion conditions. It is also helpful to break latency down into its different components (local, network, batch, consensus latency).
On the surface throughput seems simple enough: how many transactions per second can a system process? But there are two main difficulties: what exactly is a “transaction”, and are we measuring what a system does today or what it could potentially do?
While “transactions per second” (tps) is the de facto standard for measuring blockchain performance, transactions as a unit of measure are problematic. For systems that offer general programmability (smart contracts) or even limited functionality like Bitcoin’s multi-way transactions or multi-signal validation options, the fundamental problem is that:
Not all transactions are created equal.
This is clearly true in Ethereum, where transactions can include arbitrary code and modify arbitrary state. The concept of gas in Ethereum is used to quantify (and charge for) the overall amount of work a transaction is doing, but this is highly specific to the EVM execution environment. There is no easy way to compare the total amount of work done by a set of EVM transactions to a set of Solana exchanges using a BPF environment. Comparing the two to a set of Bitcoin transactions is equally concerning.
Separating the transaction layer of a blockchain into a consensus layer and an execution layer makes this more explicit. At the (pure) consensus layer, throughput can be measured in bytes added to the chain per unit of time. The execution layer is more complex.
Simpler execution layers, such as Rollup servers that only support payment transactions, avoid the difficulty of quantifying the calculation. Even in this case, payments will vary depending on the number of inputs and outputs. Payment channel transactions can vary in the number of "hops" required, which will affect throughput. Rollup server throughput depends on how well a batch of transactions can be "networked" into smaller sets of aggregated changes.
Another challenge in throughput is to go beyond empirically measuring today's performance to assess theoretical capacity. This introduces various modeling issues to assess potential capacity. First, we must identify a realistic transaction workload for the execution layer. Second, real systems almost never reach theoretical capacity, especially blockchain systems. For robustness reasons, we expect node implementations to be heterogeneous and diverse in practice (rather than all clients running a single software implementation). This makes it more difficult to accurately simulate blockchain throughput.
In summary:
Throughput claims require careful interpretation of the transaction workload and the number of validators (their number, implementations, and network connectivity). In the absence of any clear standard, historical workloads from popular networks such as Ethereum are sufficient.
Latency and throughput are often a tradeoff. This tradeoff is usually not straightforward, and latency can increase dramatically when the system load approaches maximum throughput.
Zero-knowledge rollup systems are a natural example of a throughput/latency tradeoff. Larger numbers of transactions increase proof time, and thus latency. But the on-chain footprint, both in terms of proof size and verification cost, will be spread over more transactions, larger batch sizes, and increased throughput.
End users understandably care more about the tradeoff between latency and fees than they do about latency and throughput. There is no direct reason for users to care about throughput, they just care about being able to confirm transactions quickly with the lowest possible fees (some users care more about fees, while others care more about latency). High fees are influenced by a variety of factors:
How much market demand is there to trade?
What is the overall throughput achieved by the system?
What is the total revenue provided by the system to validators or miners?
How much of this revenue is based on transaction fees or inflation rewards?
The first two factors are roughly the supply and demand curve that leads to the market clearing price (although some claim that miners acting like cartel firms drive fees above that point). All else being equal, higher throughput should lead to lower fees, but there are more factors to deal with.
Points 3 and 4 above in particular are fundamental issues for blockchain system design, but we lack good principles for both of them. We have some understanding of the benefits and drawbacks of giving miners inflationary rewards relative to transaction fees. However, despite many economic analyses of blockchain consensus protocols, we still do not have a widely accepted model for how much revenue needs to go to validators. Today, most systems are built on an educated guess about how much revenue is enough to keep validators working honestly without stifling actual use of the system. In a simplified model, it can be shown that the cost of mounting a 51% attack is proportional to the reward to the validator.
Raising the cost of attack is a good thing, but we also don’t know how much security is “enough”. Imagine you are considering going to two amusement parks. One of them claims to spend 50% less money on vehicle maintenance than the other. Is it a good idea to go to this park? It could be that they are more efficient and get equivalent security for less money. Another possibility is that it costs more than necessary to make the rides safe, without any benefit. But it could also be that the first park is dangerous. Blockchain systems are similar. Once throughput is taken out, blockchains with lower fees have lower fees because they have less rewards (and therefore incentives) for validators. We don’t have good tools right now to assess whether this is reasonable, or whether it makes the system vulnerable to attack. In summary:Comparing fees between different systems can be misleading. While transaction fees are important to users, they are affected by many factors besides the system design itself. Throughput is a better metric for analyzing the entire system.
It is hard to fairly and accurately assess performance. The same applies to measuring the performance of a car. Just like blockchain, different people care about different things. For cars, some users care about top speed or acceleration, some care about gas mileage, and some care about towing capacity. None of these are easy to get accurate values for. In the United States, for example, the Environmental Protection Agency has detailed guidelines for how gas mileage should be evaluated and how it must be presented to users at dealerships.
The blockchain space is a long way from this level of standardization. In some areas, in the future we may use standardized workloads to evaluate the throughput of a system, or standardized graphs to represent latency distributions. For now, the best approach for evaluators and builders is to collect and publish as much data as possible, and to describe the evaluation method in detail so that it can be replicated and compared with other systems.
Original link
Welcome to join the official BlockBeats community:
Telegram Subscription Group: https://t.me/theblockbeats
Telegram Discussion Group: https://t.me/BlockBeats_App
Official Twitter Account: https://twitter.com/BlockBeatsAsia