Econometrics

2,754 questions on Econometrics, part of Economics & Finance. Below are 12 of them in full, each answered in plain language.

Questions & explanations

1. Describe an example where combining lesion and neuroimaging methods improves causal inference.

Suppose researchers want to know if the anterior insula is causal for feeling disgust. They could study patients with lesions to that region and find they report less disgust to disgusting stimuli, suggesting causation. Then, using fMRI in healthy participants, they see the insula activates during disgust experiences, but this is correlational. Combining both, if the lesion patients also show no activation in connected regions during disgust, that strengthens the causal link. Additionally, they could apply TMS over the insula (if accessible) in healthy people to see if it temporarily reduces disgust sensitivity. The convergence of lesion, neuroimaging, and stimulation methods provides stronger causal evidence than any single technique alone.

2. Give an example of a panel cointegration analysis in economics.

A researcher wants to study the long-run relationship between GDP, investment, and trade openness across countries over 20 years. First, she uses panel unit root tests and finds each variable is non-stationary. Then she runs a panel regression of GDP on investment and trade, and applies the Pedroni test to the residuals. The test shows the residuals are stationary, so the variables are cointegrated. This implies a stable long-run equilibrium: GDP changes are tied to investment and trade. She then estimates the cointegrating coefficients using dynamic OLS, controlling for endogeneity. The results suggest that a 1% increase in investment raises GDP by 0.3% in the long run. This analysis helps policymakers understand growth determinants.

3. What is the multiple comparisons problem in neuroimaging and how is it addressed?

In neuroimaging, researchers test statistical significance at many brain locations (often over 100,000 voxels). By chance, some locations will appear significant even if there is no true effect. This is the multiple comparisons problem. Standard approaches control false positives by adjusting the significance threshold. The Bonferroni correction divides the alpha level (e.g., 0.05) by the number of tests, which is very strict. More advanced methods like false discovery rate (FDR) control the expected proportion of false positives among those declared significant. Cluster-based correction uses the size of contiguous activated regions. These methods reduce false positives but also reduce statistical power to detect true effects.

4. How do you test for cointegration in a panel of non-stationary variables?

To test for cointegration, we first confirm that each variable has a unit root using panel unit root tests. Then we estimate a long-run relationship, for example by regressing one variable on others. For panel cointegration tests, we examine whether the residuals from this regression are stationary. The Pedroni test is a common method; it allows for heterogeneous slopes and fixed effects. It computes seven statistics based on the residuals' autoregressive behavior. If the residuals are stationary, we conclude the variables are cointegrated. This means they move together in the long run, even if individually they wander. The next step is to estimate the cointegrating vector using methods like fully modified OLS or dynamic OLS.

5. How do you estimate a spatial lag model for panel data?

To estimate a spatial lag model, we specify the equation: Y_it = ρ * W * Y_it + X_it β + μ_i + ε_it, where W is the spatial weight matrix. ρ is the spatial autocorrelation coefficient. The term W*Y_it is the spatially lagged dependent variable, which is endogenous because neighbors' outcomes are correlated. Common estimation methods include maximum likelihood (ML) and generalized method of moments (GMM). ML requires assuming normal errors and is computationally intensive. GMM uses instruments, such as higher-order spatial lags or time lags, to handle endogeneity. Panel fixed or random effects are included to control for unit-specific factors. The model can also be estimated with bias-corrected methods for short panels.

6. What is the main idea behind Lewbel's heteroskedasticity-based identification method?

Lewbel's method creates an instrument from the model's own heteroskedasticity when no external instrument is available. It uses the fact that if the error terms have different variances across groups or with a variable, you can form an instrument by multiplying a centered variable (e.g., a regressor minus its mean) by the residuals from a first-stage regression. This constructed instrument is valid if the heteroskedasticity is correlated with the endogenous variable. The approach is useful when traditional instruments are weak or missing. The method relies on the assumption that the heteroskedasticity is not shared with the second-stage error. It is often used as a robustness check or when no other instrument exists.

7. How do you estimate bid functions in a first-price sealed-bid auction?

In a first-price auction, the highest bidder wins and pays her own bid. The bid function relates a bidder's private value to her optimal bid. To estimate it, we assume bidders are risk-neutral and have independent private values. We derive the equilibrium bid function from the model. Then we use data on bids and the number of bidders to estimate the value distribution. One common method is the 'simulation-based' approach: we simulate bids for candidate value distributions and choose the one that best matches actual bids. Another method is the 'parametric' approach, where we assume a functional form for values and use maximum likelihood. The key is that bids are monotonic in values, so we can invert the relationship.

8. What are spatial lag and spatial error models for panel data?

Spatial panel models extend standard panel models to account for spatial dependence between units. A spatial lag model includes a term for the weighted average of the dependent variable from neighboring units. This captures spillover effects: one unit's outcome affects others. A spatial error model puts the spatial dependence in the error term, modeling unobserved shocks that spread across space. Both models use a spatial weight matrix that defines which units are neighbors. Estimation typically uses maximum likelihood or instrumental variables to avoid endogeneity. Panel data add fixed or random effects to control for time-invariant heterogeneity. These models are common in regional science and public economics.

9. What is a lesion study in neuroscience and what causal question does it answer?

A lesion study examines individuals who have damage (a lesion) to a specific brain region, often due to stroke, injury, or surgery. By comparing the behavior or cognitive abilities of patients with a lesion to those without, researchers infer that the damaged region is necessary for the function that is impaired. For example, patients with damage to Broca's area often have difficulty speaking fluently, suggesting that area is causal for language production. Lesion studies provide strong causal evidence because the damage is usually not under the experimenter's control but is naturally occurring. However, lesions are rarely confined to a single area, and the brain can reorganize, complicating interpretation.

10. Give an example where double selection would be useful in economics.

Suppose an economist wants to estimate the effect of a job training program on wages. There are many possible control variables: age, education, previous work experience, location, marital status, many past test scores, and dozens of other characteristics. With only a few thousand workers, it is impossible to include all these variables in a regression without overfitting. Double selection would first use Lasso to pick which variables best predict wages (the outcome) and which best predict program participation (the treatment). The union of these variables is then included in a final regression. This yields a reliable estimate of the training effect while controlling for the most relevant confounders.

11. What is a fixed effects model and what problem does it solve?

A fixed effects model is a way to analyze panel data (data on the same units over time) that controls for unobserved factors that do not change over time, like a firm's management quality. It removes the bias from these constant unobserved factors by using only the variation within each unit over time. For example, if you study how training affects worker wages, fixed effects can control for each worker's unchanging ability. This model compares each unit to itself at different times, so time-invariant differences between units do not matter. It is useful when you think those constant unmeasured factors might be related to your main variable. The model essentially uses each unit as its own control.

12. How does structural estimation differ between first-price and second-price auctions?

In a second-price auction, the highest bidder wins but pays the second-highest bid. The optimal strategy is to bid your true private value, so the bid function is simply identity. This makes structural estimation straightforward: we observe bids and take them as values directly. In contrast, first-price auctions require solving a differential equation to link bids to values, because bidders shade their bids. Estimation for first-price is more complex, often needing numerical methods or simulation. Second-price auctions provide a direct observation of values, but in practice they are less common due to strategic concerns. Both models can be extended to include risk aversion or correlated values.

More Economics & Finance topics

This page shows 12 of 2,754 questions on this topic. The full set, with progress tracking and five agent perspectives per question, is in the JupiteX app — browse the exam catalogue or browse the Learn library.