The Congressional scorekeepers—the Congressional Budget Office (CBO) and the Joint Committee on Taxation (JCT)—have an important job. At their best, they serve an important role in budget restraint, shining light on the budget and economic effects of Congress’s spendthrift policies. Members of Congress, however, have been increasingly dismissive of the scorekeepers’ work.
Throughout our series (and in a recent WSJ op-ed by Josh and Ben), we’ve examined why trust in the scorekeepers has eroded:
Despite nods to increased transparency, their models and methods remain opaque. This is particularly true within the JCT.
Their models are updated with little documentation or warning, sometimes shortly after the scorekeepers have published politically contentious results.
Their estimates and reports often convey more confidence in projections than the evidence warrants.
At a recent House Budget hearing, Chair Jodey Arrington, paraphrasing Churchill, reminded Director Swagel that “the price of greatness is accountability.” Without it, members will increasingly disregard the scorekeepers’ work and legislate without fiscal restraint. But rebuilding that trust will require overhauling how the CBO and JCT operate. Congress can’t expect omniscience, but they should demand full transparency and accountability.
Reform #1: Demand Replicability
Congress should require that the scorekeepers publish their models. Reports and scores should be accompanied by code, data, and documentation to allow for the replication of all results and exploration of underlying assumptions.1 When internal improvements are made to long-lasting models, the changes should be posted with accompanying documentation and explanations for the changes.
This demand isn’t new. The CBO Show Your Work Act has been introduced regularly in Congress since 2017 by Congressman Warren Davidson and Senator Mike Lee. It would require the CBO “to make available to Congress and the public each fiscal model, policy model, and data preparation routine that the CBO uses to estimate the costs and other fiscal, social, or economic effects of legislation.”2
This is already the industry standard in economics. Major economics journals require authors to post replication packages that allow for the reproduction of a paper’s results. The concept is also not entirely novel to the CBO. The agency already encourages its staff “to post data, computer code, and documentation associated with models and projects” to its GitHub page, but only “to the extent practicable.” They must have a narrow view of what is practicable. In 2023, the CBO had published four of its models to its GitHub page. In the two years since, it has added just two models to its repository.3
But a replicability standard for new reports and scores would go far beyond the CBO’s meager transparency efforts. It would be entirely new to the JCT. The scorekeepers make several arguments against such radical transparency:
The Cost: The most common argument against a replicability standard is that it would be too costly for an agency with a budget of about $70 million and a staff of under 300. These objections may have some merit—although they don’t explain why the CBO has released only six of its models in the last five years.4 At the recent House Budget hearing, CBO Director Phillip Swagel contended that the lack of recent GitHub uploads was due to limited resources, noting that his staff had been focused on producing analyses and scores for the One Big, Beautiful Bill Act.5 To reduce the expected costs, lawmakers could initially limit the replicability requirement to regularly published reports (e.g., the CBO’s 10-year and long-term budget and economic outlooks, JCT’s annual report on tax expenditures) and legislation expected to have a meaningful effect on the budget or the economy. There is precedent for this approach. House Rule XIII requires the scorekeepers to perform a dynamic analysis on any legislation that is projected to have a gross budget effect of more than 0.25 percent of GDP in any year over the 10-year budget window. Few bills meet this standard, but lawmakers could set a lower GDP threshold for the replicability requirement (perhaps 0.1 percent of GDP), which could be lowered over time as more models are successfully posted.
Confidentiality: A separate concern is that the scorekeepers’ models use confidential data that is legally protected from disclosure. For example, the scorekeepers occasionally rely on highly sensitive taxpayer data or personal health claims data to build their scores. The CBO Show Your Work Act offers a reasonable mechanism to address this. When confidential data is employed, the scorekeepers should still provide a list of all variables used, descriptive statistics of each variable,6 and a reference to the underlying statute that prevents disclosure. As Matt Jensen argued in a 2018 National Affairs piece, the scorekeepers should provide sufficient information to “allow outside replicators to construct a test dataset that enables them to run the data-preparation routines and model code from beginning to end.”
Gaming the Score: Some worry that full transparency would allow lawmakers to game the score—writing legislation to earn the best possible estimate even if it produces suboptimal policy. But that already happens, just in a more costly and inefficient way. Today, congressional staff will often ask the scorekeepers to rerun estimates after making small tweaks to legislative language in hopes of a better score. This iterative process is frustrating and time consuming to staffers and the scorekeepers. It would happen less often if there was a better understanding of the nuances of the methodology. This would reduce the demands on scorekeepers, potentially lowering their long-term costs. More importantly, replicability could improve the models to reduce the likelihood that suboptimal legislation scores well.
The big question is how the scorekeepers can meet the replicability standard. To achieve that, we think the requirements must be paired with other reforms to the agencies. We discuss these below.
Reform #2: Ombudsman
House Budget Chair Jodey Arrington has asked for an external audit of the CBO. At the recent hearing, Director Swagel said he would be willing to explore such an idea. While a good start, a one-time exploration of the CBO’s models and processes is unlikely to deliver the radical transparency needed to restore trust in the agency. Instead, Congress should appoint a permanent ombudsman to streamline the transparency efforts of both agencies.
The ombudsman would be Congress’s transparency advocate within the agencies. The office of the ombudsman would assist the CBO and the JCT in ensuring that the replicability requirements are being met. This would include ensuring that all relevant code and underlying data are released. In the case of confidential data, the ombudsman would certify that the agencies are withholding only legally protected data and ensure that all required variable descriptions and summary statistics are published.
Beyond their role in meeting the replicability requirements, the ombudsman would provide regular reports to Congress on the transparency efforts of both agencies. The reports would identify the obstacles to increased transparency, provide assessments on whether the threshold for replicability should be changed, and outline potential further reforms. The ombudsman could also serve as an important intermediary between Congressional staff and the scorekeepers when internal or external concerns about models are raised. Over time, the ombudsman could also serve as an external check on key models and projections.7
The existing leadership of the CBO and JCT would continue to be accountable to Congress. Their roles wouldn’t change.8 They are already tasked with considerable duties including managing hundreds of employees and thousands of member requests. Given their many responsibilities, tight deadlines, and budget constraints, it is no wonder that transparency efforts have lagged. The office of the ombudsman would ensure the transparency efforts can be advanced without burdening existing leadership with more work.
Reform #3: Humility, Sensitivity Analyses and Confidence Ratings
To restore trust, the scorekeepers must exercise more humility when reporting certain estimates to Congress. As we noted in an earlier essay, there are times when Congress needs a number to fulfill certain budgeting rules. But often, the most highly publicized number in a projection is peripheral to the score. In these cases, a number isn’t needed. For example, when Congress considers legislation affecting health laws, the CBO will estimate the legislation’s effect on the number of uninsured. This inevitably leads to headlines that cite the CBO’s point estimate, typically rounded to the nearest 100,000. But that level of precision is rarely warranted. The faux precision undermines trust as these numbers and the underlying model inevitably change, sometimes within months after the headline. To improve trust, Congress should demand that the scorekeepers rethink this approach.
Robust Sensitivity Analysis: The scorekeepers should regularly produce sensitivity analyses for major legislation and reports to convey the inherent uncertainty of their estimates. The CBO already does this for a handful of economic and budget assumptions in the annual budget projections, but this is the exception not the rule. And even these estimates are of limited value. They are generally limited to a partial analysis that only explores the effects of changing one parameter at a time. A more rigorous approach would estimate the joint distribution of these parameters to produce meaningful confidence intervals. This would be quite the technical lift for the scorekeepers, which is another reason why a replicability standard is so important. While the scorekeepers may not always have the bandwidth for robust sensitivity analysis, independent researchers could advance these efforts if they have access to the CBO’s underlying models and data.
Confidence Ratings: The CBO and the JCT should immediately develop a systematic ranking of their confidence in their estimates. The U.S. Intelligence Community already uses “analytic confidence” ratings to assess how confident they are in their intelligence assessment. The CBO and the JCT, with the input of the ombudsman, would assess their confidence in major components of scores and analyses. Low confidence might be assigned to a policy being scored for the first time, if the data used for the score are limited, or if the model is being actively updated. Medium confidence might be assigned when the score is based on standard models, but noisy data. Lastly, high confidence might be assigned when that result rests on standard models, clean data, and broad agreement in the academic and policy literature.
The office of the ombudsman would play an important role in these efforts. First, it could help develop the rating system. Second, it could serve as an external validator that examines, post hoc, whether the CBO and JCT are appropriately rating their scores.
Reform #4: Narrowing the Mission
As we noted, the CBO and JCT are under considerable strain. They must answer thousands of member requests to score specific legislation, create regular reports, and maintain many economic and budget models. Beyond these asks, they are regularly asked to opine on large policy issues. In just the last year, the CBO has studied the budget consequences of weight-loss drugs and the impact of recent immigration trends on state budgets. All with a relatively small budget and staff.
In short, we are asking too much of them.
Congress should streamline the scorekeepers’ work to focus on issues directly related to the federal budget. Congress should limit the scorekeepers’ work to:
1. Legislative scores
2. Reports on the short-term and long-term U.S. budget and economy
3. Economic and policy analyses expected to have immediate utility for developing these scores and reports. This would include analysis on the effect changes in debt have on interest rates and behavioral effects of tax policy. It could also include analyses such as the budget consequences of weight-loss medicines.
4. Matters of congressional concerns that the CBO and the JCT are uniquely qualified to answer.
Other analytical questions may be better served by external researchers, the Government Accountability Office, the Congressional Research Service, or executive agencies. The CBO and the JCT have two competitive advantages over external researchers. First, they have access to confidential government data (e.g., IRS data, Medicare and Medicaid claims data). Second, they can speak directly with regulators and agency officials about how different policies would be enacted. The CBO and the JCT should not be asked to opine on research questions that do not require either of these or are not directly related to legislative scoring or budget analyses.
For example, the CBO recently released a report on the long-term economic effects of climate change. This meta-analysis summarized several external estimates on the effect of GDP from climate change one century from now. Ironically, it included sensitivity analyses that are lacking from most of CBO’s scores and budget analyses. Given that CBO’s long-term budget modeling only extends 30 years into the future, it is unlikely that this analysis would inform any of CBO’s existing reports or legislative scores. Nor was the CBO uniquely positioned to perform this analysis. In short, the analysis could have just as easily been performed by outside research groups rather than overworked scorekeepers that argue they have insufficient resources for basic transparency efforts.
Conclusion
Trillion-dollar deficits and unprecedented debt levels mean we need the scorekeepers more than ever. They remain essential tools for improving future budgets, but only if their analyses command trust. Those concerned about the nation’s fiscal trajectory should insist that the agencies’ work withstand all scrutiny. The scorekeepers should be positioned as credible voices for restraint, not controversy. A commitment to real transparency and accountability would help them serve that role.
Specifically, the CBO and the JCT should release the unprocessed data, any code used to prepare the data, and the code used in their final analysis.
Also, see Matt Jensen’s excellent 2018 National Affairs piece on the subject.
The CBO also publishes interactive Excel workbooks that allow external audiences to explore how certain economic and policy assumptions would affect CBO’s score. While these are useful for researchers to understand CBO’s thinking, they fall far too short of the replicability standard that lawmakers should demand.
The long-term costs are less obvious. While there would be an upfront cost, replicability could eventually reduce demands on the scorekeepers. A replicability standard would require a commitment to best practices in model development. This would pay dividends in the future as CBO and JCT staff are not forced to rewrite or decipher preexisting models written by departed colleagues.
This explanation does not explain why the CBO failed to upload more models last year.
Descriptive statistics would include “averages, standard deviations, number of observations, and correlations to other variables.”
Currently, the CBO will occasionally review past projections for accuracy, but this is generally on an ad hoc basis.
The Senate and House Budget Committees have jurisdiction over the CBO. The chair of the JCT is currently House Ways and Means Chairman Jason Smith and the Vice Chair is Senate Finance Committee Chairman Mike Crapo. Smith and Crapo will switch roles at the start of the second session (January 2026).



