Research
Publications
-
Structured Payment in Pawnshop Borrowing: Mandates vs. Choice (with Craig McIntosh, Isaac Meza, Joyce Sadka, and Enrique Seira), Review of Economic Studies, Forthcoming
Pawn loans offer borrowers a substantial degree of repayment flexibility in exchange for a harsh penalty in case of default: forfeit of collateral worth more than the loan amount along with any payments made toward recovery. Using a large RCT conducted in Mexico City, we document key stylized facts about pawn lending and explore the merits of replacing flexibility with structured repayment contracts in this important but understudied form of credit. Our experimental design includes a mandatory frequent-payments arm, a (status quo) flexible payments arm, and a choice between the two. This design point-identifies not only the average treatment effect, but also the effects of treatment on the treated and the untreated along with the average selection on gains, allowing a rigorous study of mandates versus choice. Although the average treatment effect of assigning borrowers to structured payments is a 19% decrease in their financial cost and a 17.5% decrease in the probability of default, only 11% of borrowers choose structured repayment contracts voluntarily. We show that structured repayment reduces financial costs for nearly all borrowers, including those who would not freely choose it, and find no evidence of selection on gains in cost savings.
-
Efficient Semiparametric Estimation of Marginal Treatment Effects with Genetic Instrumental Variables (with Ashish Patel and Stephen Burgess), Annals of Applied Statistics, Forthcoming
code
Alcohol misuse is a key target of public health strategies aimed at reducing cardiovascular risk. The effect of excessive alcohol consumption on blood pressure may vary systematically with individuals’ unobserved propensity to engage in heavy drinking, complicating causal inference with observational data. The marginal treatment effects framework uses an instrumental variable for treatment choice (excessive alcohol consumption) to study how selection into treatment is linked with the treatment effect. We explore the use of a genetic instrument within this framework, which is challenging because genetic compliers (individuals for whom a change in the instrument changes their treatment choice) are likely to be a small proportion of the overall sample. This can lead to greater sampling uncertainty in the tails of the propensity score distribution, i.e., the conditional probability of choosing treatment, and in turn poor estimation of causal estimands that measure heterogeneous treatment effects. We show that the use of efficient influence functions of target estimands improves estimation in terms of robustness to sampling uncertainty in nonparametrically estimated propensity scores. We find evidence of reverse selection on gains: individuals most prone to excessive alcohol consumption experience larger adverse effects on blood pressure.
-
Environmental Lead Risk in the 21st Century (with Mengli Chen, Ludovica Gazze, Reshmi Das, Jerome Nriagu, Yigal Erel, Edward Boyle, Caroline Taylor, and Dominik Weiss), Communications Earth & Environment, 2025, 6:776
code
Lead has been central to technological development for centuries; however, its release into the environment and subsequent human exposure pose significant public health risks. The review presented here critically assesses the contemporary environmental lead risk as global lead production and use are rapidly increasing, largely driven by the rising demand for electrification. We show that environmental lead exposure persists today due to legacy contamination, ongoing coal usage, and insufficient protection of workforces during production, use, and recycling of lead-acid batteries and other lead-containing products, particularly in low- and middle- income countries. We estimate that contemporary childhood lead exposure alone leads to an annual global economic loss exceeding $3.4 trillion (2021 US dollars adjusted for purchasing power parity), with pronounced disparities between high- and low- and middle- income countries. To prevent a large-scale resurgence in lead exposure, we identify four critical areas for urgent policy intervention.
-
Selecting Invalid Instruments to Improve Mendelian Randomization with Two-Sample Summary Data (with Ashish Patel, Verena Zuber, and Stephen Burgess), Annals of Applied Statistics, 2024, 18 (2), pp. 1729–1749
pre-print
Mendelian randomization (MR) is a widely-used method to estimate the causal relationship between a risk factor and disease. A fundamental part of any MR analysis is to choose appropriate genetic variants as instrumental variables. Genome-wide association studies often reveal that hundreds of genetic variants may be robustly associated with a risk factor, but in some situations investigators may have greater confidence in the instrument validity of only a smaller subset of variants. Nevertheless, the use of additional instruments may be optimal from the perspective of mean squared error even if they are slightly invalid; a small bias in estimation may be a price worth paying for a larger reduction in variance. For this purpose, we consider a method for "focused" instrument selection whereby genetic variants are selected to minimise the estimated asymptotic mean squared error of causal effect estimates. In a setting of many weak and locally invalid instruments, we propose a novel strategy to construct confidence intervals for post-selection focused estimators that guards against the worst case loss in asymptotic coverage. In empirical applications to: (i) validate lipid drug targets; and (ii) investigate vitamin D effects on a wide range of outcomes, our findings suggest that the optimal selection of instruments does not involve only a small number of biologically-justified instruments, but also many potentially invalid instruments.
-
Identifying Causal Effects in Experiments with Spillovers and Non-compliance (with Camilo García-Jimeno, Rossa O’Keeffe-O’Donovan, and Alejandro Sánchez-Becerra), Journal of Econometrics, 2023, 235 (2), pp. 1589–1624
pre-printslides
This paper shows how to use a randomized saturation experimental design to identify and estimate causal effects in the presence of spillovers—one person’s treatment may affect another’s outcome—and one-sided non-compliance—subjects can only be offered treatment, not compelled to take it up. Two distinct causal effects are of interest in this setting: direct effects quantify how a person’s own treatment changes her outcome, while indirect effects quantify how her peers’ treatments change her outcome. We consider the case in which spillovers occur within known groups, and take-up decisions are invariant to peers’ realized offers. In this setting we point identify the effects of treatment-on-the-treated, both direct and indirect, in a flexible random coefficients model that allows for heterogeneous treatment effects and endogenous selection into treatment. We go on to propose a feasible estimator that is consistent and asymptotically normal as the number and size of groups increases. We apply our estimator to data from a large-scale job placement services experiment, and find negative indirect treatment effects on the likelihood of employment for those willing to take up the program. These negative spillovers are offset by positive direct treatment effects from own take-up.
-
A Framework for Eliciting, Incorporating, and Disciplining Identification Beliefs in Linear Models (with Camilo García-Jimeno), Journal of Business & Economic Statistics, 2021, 39 (4), pp. 1038–1053
pre-print
To estimate causal effects from observational data, an applied researcher must impose beliefs. The instrumental variables exclusion restriction, for example, represents the belief that the instrument has no direct effect on the outcome of interest. Yet beliefs about instrument validity do not exist in isolation. Applied researchers often discuss the likely direction of selection and the potential for measurement error in their articles but lack formal tools for incorporating this information into their analyses. Failing to use all relevant information not only leaves money on the table; it runs the risk of leading to a contradiction in which one holds mutually incompatible beliefs about the problem at hand. To address these issues, we first characterize the joint restrictions relating instrument invalidity, treatment endogeneity, and non-differential measurement error in a workhorse linear model, showing how beliefs over these three dimensions are mutually constrained by each other and the data. Using this information, we propose a Bayesian framework to help researchers elicit their beliefs, incorporate them into estimation, and ensure their mutual coherence. We conclude by illustrating our framework in a number of examples drawn from the empirical microeconomics literature.
-
Identifying the Effect of a Mis-classified, Binary, Endogenous Regressor (with Camilo García-Jimeno), Journal of Econometrics, 2019, 209 (2), pp. 376–390
pre-printNBER version
This paper studies identification of the effect of a mis-classified, binary, endogenous regressor when a discrete-valued instrumental variable is available. We begin by showing that the only existing point identification result for this model is incorrect. We go on to derive the sharp identified set under mean independence assumptions for the instrument and measurement error. The resulting bounds are novel and informative, but fail to point identify the effect of interest. This motivates us to consider alternative and slightly stronger assumptions: we show that adding second and third moment independence assumptions suffices to identify the model.
-
A Generalized Focused Information Criterion for GMM (with Minsu Chang), Journal of Applied Econometrics, 2018, 33 (3), pp. 378–397
pre-printcode
This paper proposes a criterion for simultaneous generalized method of moments model and moment selection: the generalized focused information criterion (GFIC). Rather than attempting to identify the “true” specification, the GFIC chooses from a set of potentially misspecified moment conditions and parameter restrictions to minimize the mean squared error (MSE) of a user‐specified target parameter. The intent of the GFIC is to formalize a situation common in applied practice. An applied researcher begins with a set of fairly weak “baseline” assumptions, assumed to be correct, and must decide whether to impose any of a number of stronger, more controversial “suspect” assumptions that yield parameter restrictions, additional moment conditions, or both. Provided that the baseline assumptions identify the model, we show how to construct an asymptotically unbiased estimator of the asymptotic MSE to select over these suspect assumptions: the GFIC. We go on to provide results for postselection inference and model averaging that can be applied both to the GFIC and various alternative selection criteria. To illustrate how our criterion can be used in practice, we specialize the GFIC to the problem of selecting over exogeneity assumptions and lag lengths in a dynamic panel model, and show that it performs well in simulations. We conclude by applying the GFIC to a dynamic panel data model for the price elasticity of cigarette demand.
-
Using Invalid Instruments on Purpose: Focused Moment Selection and Averaging for GMM, Journal of Econometrics, 2016, 195 (2), pp. 187–208
pre-printappendixcode
In finite samples, the use of a slightly endogenous but highly relevant instrument can reduce mean-squared error (MSE). Building on this observation, I propose a novel moment selection procedure for GMM—the Focused Moment Selection Criterion (FMSC)—in which moment conditions are chosen not based on their validity but on the MSE of their associated estimator of a user-specified target parameter. The FMSC mimics the situation faced by an applied researcher who begins with a set of relatively mild “baseline” assumptions and must decide whether to impose any of a collection of stronger but more controversial “suspect” assumptions. When the (correctly specified) baseline moment conditions identify the model, the FMSC provides an asymptotically unbiased estimator of asymptotic MSE, allowing us to select over the suspect moment conditions. I go on to show how the framework used to derive the FMSC can address the problem of inference post-moment selection. Treating post-selection estimators as a special case of moment-averaging, in which estimators based on different moment sets are given data-dependent weights, I propose simulation-based procedures for inference that can be applied to a variety of formal and informal moment-selection and averaging procedures. Both the FMSC and confidence interval procedures perform well in simulations. I conclude with an empirical example examining the effect of instrument selection on the estimated relationship between malaria and income per capita.
-
Portfolio Selection: An Extreme Value Approach (with Jeffrey Gerlach), Journal of Banking & Finance, 2013, 37 (2), pp. 305–323
pre-print
We show theoretically that lower tail dependence (χ), a measure of the probability that a portfolio will suffer large losses given that the market does, contains important information for risk-averse investors. We then estimate χ for a sample of DJIA stocks and show that it differs systematically from other risk measures including variance, semi-variance, skewness, kurtosis, beta, and coskewness. In out-of-sample tests, portfolios constructed to have low values of χ outperform the market index, the mean return of the stocks in our sample, and portfolios with high values of χ. Our results indicate that χ is conceptually important for risk-averse investors, differs substantially from other risk measures, and provides useful information for portfolio selection.
-
Measuring Altruism in a Public Goods Experiment: A Comparison of U.S. and Czech Subjects (with Lisa Anderson and Jeffrey Gerlach), Experimental Economics, 2011, 14 (3), pp. 426–437
This paper compares contributions to an experimental public good across the United States and Czech Republic, using a design that allows us to distinguish between altruism and decision error. Czech subjects contribute significantly more than American subjects, and further analysis reveals that this result cannot be attributed to the confounding effects of gender or decision error. Instead, preferences for altruism appear to differ across groups: Czechs are more altruistic than Americans and men are more altruistic than women.
-
Yes, Wall Street, There is a January Effect! Evidence from Laboratory Auctions (with Lisa Anderson and Jeffrey Gerlach), Journal of Behavioral Finance, 2007, 8 (1), pp. 1–8
There is a large literature using financial market data on the causes of a “January effect,” which produces higher stock prices in January than in other months of the year. We present the first experimental study of this phenomenon in the context of two well-known auction experiments. After controlling for variables that could influence subject bids, such as differences in private values, cumulative earnings, and learning effects, the prices in the January markets were systematically higher than those in December, a difference that is economically large and statistically significant. The results provide support for the conjecture that psychological factors may contribute to the well-documented January effect in empirical stock market data.
Working Papers
-
Bayesian Double Machine Learning for Causal Inference (with Laura Liu)
arXivslides
This paper proposes a simple, novel, and fully-Bayesian approach for causal inference in partially linear models with high-dimensional control variables. Off-the-shelf machine learning methods can introduce biases in the causal parameter known as regularization-induced confounding. To address this, we propose a Bayesian Double Machine Learning (BDML) method, which modifies a standard Bayesian multivariate regression model and recovers the causal effect of interest from the reduced-form covariance matrix. Our BDML is related to the burgeoning frequentist literature on DML while addressing its limitations in finite-sample inference. Moreover, the BDML is based on a fully generative probability model in the DML context, adhering to the likelihood principle. We show that in high dimensional setups the naive estimator implicitly assumes no selection on observables—unlike our BDML. The BDML exhibits lower asymptotic bias and achieves asymptotic normality and semiparametric efficiency as established by a Bernstein-von Mises theorem, thereby ensuring robustness to misspecification. In simulations, our BDML achieves lower RMSE, better frequentist coverage, and shorter confidence interval width than alternatives from the literature, both Bayesian and frequentist.
-
Experimenting with Spillovers: A Guide for Practitioners (with Alejandro Sánchez-Becerra), Invited Chapter, Oxford Handbook of Impact Evaluation (email me for a copy)
A fundamental assumption in the introductory textbook approach to randomized controlled trials is the absence of spillovers: that assigning a particular treatment to one participant has no causal effect on any other participant. In practice, however, spillovers between experimental participants are common in social science, arising through mechanisms such as displacement and contagion. This chapter presents a comprehensive practitioner’s guide to designing and analyzing randomized controlled trials in the presence of spillovers. We discuss four main approaches: (1) randomized saturation designs that decompose direct and spillover effects by varying treatment intensity across clusters; (2) cluster designs that exploit natural variation in program eligibility to estimate spillovers; (3) group formation experiments that study peer effects by randomly grouping participants; and (4) network targeting experiments that test strategies for leveraging spillovers to improve program effectiveness. For each approach, we detail the necessary assumptions and the meaning of the causal effects that can be recovered. We also provide practical implementation guidance and discuss a variety of real-world applications, ranging from labor market interventions to education and public health. The chapter concludes with a brief overview of extensions, including non-compliance, exposure mappings, and within-cluster network analysis.
In Progress
- Inference for the 21st Century (with Frank Schorfheide)
- Bayesian Causal Inference in High Dimensions (with Laura Liu)
- Evaluating Robustness at Scale (with Johanna Einsiedler and Elizabeth Rhodes)
- ECLIPS: the Elevated Childhood Lead Interagency Prevalence Study (with Ludovica Gazze et al.)
- To Link or Not to Link? Estimating Long-run Treatment Effects from Historical Data (with Camilo García-Jimeno and Ezra Karger)
- A History of Violence: Forced Displacement and de facto Land Reform in Rural Colombia (with Camilo García-Jimeno)