Modern optimization in observational studies
"Phân tích và tối ưu hóa phương pháp nghiên cứu quan sát bằng các kỹ thuật hiện đại để cải thiện độ chính xác và hiệu suất của kết quả."
Statistics and Probability
Luan An
Luận án
Năm xuất bản
Số trang
177
Thời gian đọc
27 phút
Lượt xem
0
Lượt tải
0
Phí lưu trữ
50 Point
Tổng quan nhanh
- Chủ đề:
- Optimizing Observational Studies for Causal Inference
- Số trang:
- 177 trang
- Trường:
- university of pennsylvania
- Chuyên ngành:
- Statistics and Probability
- Tác giả:
- Colin Burton Fogarty
- Năm:
- 2016
Tóm tắt nội dung luận án
I.Optimizing Observational Studies for Causal Inference
Modern optimization techniques are crucial in observational studies. They address inherent challenges for valid causal inference. Historically, matching algorithms gained popularity. These methods pair treated units with similar control units. The goal is to adjust for overt biases. Computational advances in network flow optimization significantly boosted the use of matching. This field explores how contemporary optimization approaches solve diverse problems. Beyond traditional matching, new methodologies enhance causal inference. Focus remains on robust estimation and bias reduction techniques. The aim is to strengthen conclusions drawn from observational data. Effective strategies minimize confounding bias. The work advances statistical rigor in real-world research.
1.1. Role of Matching in Bias Reduction Techniques
Matching serves as a fundamental bias reduction technique in observational studies. It creates balanced comparison groups. Treated individuals are matched with control individuals possessing similar covariate profiles. This process mimics randomization, a cornerstone of experimental design. The objective is to ensure comparability between groups. This comparability reduces confounding bias, a major threat to causal inference. Various matching algorithms exist. Each algorithm seeks to optimize similarity across key variables. Propensity Score Matching (PSM) is a well-known example. PSM uses a single score to summarize covariates. This simplifies the matching process. Despite its advantages, PSM requires careful implementation. It addresses observed confounding but not unmeasured confounding. The effectiveness of matching relies on accurate covariate measurement. It also depends on the absence of selection bias. Modern optimization enhances matching efficiency. It improves the quality of matched sets. This leads to more reliable causal estimates.
1.2. Computational Advances Driving Modern Optimization
Computational advances underpin the widespread adoption of modern optimization in observational studies. Complex algorithms, once impractical, are now feasible. Network flow optimization, for instance, revolutionized matching algorithms. It allows for efficient creation of optimal matched sets. This reduces the computational burden for researchers. High-performance computing enables larger datasets to be analyzed. This supports more intricate models for causal inference. Machine learning for causal inference benefits greatly from these advancements. Machine learning algorithms can identify complex relationships between variables. They help in constructing more precise propensity scores. They also assist in developing robust estimation strategies. These tools improve the ability to handle high-dimensional data. They lead to more sophisticated bias reduction techniques. The synergy between statistical theory and computational power propels the field forward. It facilitates development of methods like Targeted Maximum Likelihood Estimation (TMLE) and G-Computation. These methods offer increased robustness and efficiency. They provide a foundation for advanced sensitivity analyses.
II.Addressing Covariate Overlap for Valid Causal Inference
Covariate overlap is a critical concern in observational studies for valid causal inference. Lack of overlap occurs when treatment and control groups have distinct covariate distributions. In such scenarios, direct comparisons become unreliable. Extrapolation outside the common support can lead to biased results. Modern optimization offers solutions to define an interpretable study population. This ensures inference is conducted within relevant data ranges. It prevents drawing conclusions based on dissimilar individuals. The process identifies a subpopulation where treated and control units are sufficiently comparable. This enhances the generalizability and credibility of findings. Bias reduction techniques are applied proactively. They focus on establishing a common support region. This methodological approach strengthens the foundation for causal effect estimation.
2.1. The Maximal Box Problem for Study Population Definition
The maximal box problem provides a framework for defining an interpretable study population. It addresses the issue of covariate overlap directly. The problem seeks to identify the largest rectangular region in the covariate space. Within this region, both treated and control units are present. This ensures a sufficient density of data points for comparison. The maximal box explicitly defines the boundaries for valid inference. It prevents extrapolation to regions where treatment assignment is deterministic. This is crucial for avoiding selection bias. By delineating a common support region, researchers ensure robustness. Causal inference can proceed without assuming effects in areas lacking empirical evidence. This method enhances transparency regarding the population under study. It improves the credibility of subsequent analyses. This innovative application of optimization is a powerful bias reduction technique. It supports more responsible interpretation of causal effects in observational studies.
2.2. Avoiding Extrapolation in Causal Inference Studies
Avoiding extrapolation is paramount for sound causal inference. Extrapolation occurs when estimating treatment effects for individuals outside the common support. This means applying findings to covariate profiles not observed in both treatment and control groups. Such practices introduce significant uncertainty and potential bias. The maximal box problem directly mitigates this risk. It rigorously defines the population where comparisons are empirically supported. This prevents researchers from making unsupported claims. The method ensures that causal effect estimates reflect actual data. It fosters greater confidence in the study's conclusions. By restricting inference to overlapping covariate regions, bias reduction is achieved. This methodological discipline guards against speculative generalizations. It reinforces the principle that observational studies require careful boundary conditions for validity. This strategy is essential for robust causal inference.
III.Integer Programming for Binary Outcomes in Matched Designs
Integer programming offers a sophisticated approach for analyzing binary outcomes in matched observational studies. This optimization technique is particularly valuable for complex designs. It moves beyond simple comparisons within matched sets. Integer programming facilitates precise causal inference. It constructs confidence intervals for meaningful causal estimands. This method rigorously accounts for the intricacies of matched data structures. It provides a robust framework for statistical analysis. Furthermore, it allows for comprehensive sensitivity analyses. These analyses assess the robustness of findings to potential unmeasured confounding. The application of integer programming enhances the credibility of results. It provides a powerful tool for researchers. This approach strengthens the overall methodology for observational studies. It offers advanced bias reduction techniques.
3.1. Inference and Confidence Intervals for Causal Estimands
Integer programming enables advanced inference for causal estimands with binary outcomes. It allows for exact inference in matched observational studies. This is a significant advantage over asymptotic approximations. The method constructs exact confidence intervals for treatment effects. This provides precise quantification of uncertainty. It directly addresses the challenges posed by matched designs. Causal inference is conducted with high statistical rigor. The approach can handle various estimands. These include differences in proportions or odds ratios. The exact nature of these intervals is crucial. It ensures robust conclusions, especially with smaller sample sizes within matched sets. This method represents a powerful application of optimization. It strengthens the reliability of causal effect estimates. It ensures transparency in statistical reporting. This is a key development in bias reduction techniques.
3.2. Sensitivity Analyses for Robustness to Confounding Bias
Integer programming also facilitates robust sensitivity analyses for binary outcomes. These analyses are essential for assessing the impact of unmeasured confounding bias. Unmeasured confounders can invalidate causal inference if not addressed. The method allows researchers to quantify how strong an unmeasured confounder must be. It assesses its ability to explain away an observed treatment effect. This provides a clear measure of robustness. For matched observational studies, this capability is invaluable. It helps researchers understand the limits of their findings. The technique allows for systematic exploration of hypothetical scenarios. These scenarios involve varying degrees of unmeasured confounding. This transparency builds confidence in the study's conclusions. It is a critical component of responsible causal inference. This advanced optimization technique offers crucial insights into potential biases.
IV.Convex Optimization Sensitivity for Multiple Outcomes
Convex optimization offers an innovative approach for sensitivity analysis when multiple outcome variables are present. In many observational studies, researchers investigate several endpoints simultaneously. This creates a challenge for statistical power and interpretation. Accounting for multiple comparisons can lead to a loss of statistical power. This makes it harder to detect true effects. Convex optimization provides a method to attenuate this power loss. It allows for a more efficient assessment of findings' robustness. This applies particularly to the presence of unmeasured confounding. The technique offers a structured way to evaluate the impact of potential biases across different outcomes. This integrated approach strengthens causal inference. It provides a comprehensive framework for bias reduction techniques. It ensures a more holistic understanding of treatment effects.
4.1. Attenuating Power Loss with Multiple Comparisons
Accounting for multiple comparisons is standard practice in statistical analysis. However, it often reduces the power to detect significant effects. Convex optimization provides a strategy to attenuate this power loss. It does so while still addressing the issue of multiple outcome variables. The method intelligently combines information across outcomes. This avoids overly conservative adjustments. Researchers can maintain a reasonable level of statistical power. This is crucial for discovering true treatment effects. The technique allows for more nuanced interpretations of findings. It prevents discarding important associations due to stringent multiplicity corrections. This application of optimization makes studies more efficient. It ensures a balanced approach between Type I error control and Type II error avoidance. This contributes to more effective causal inference.
4.2. Assessing Robustness to Unmeasured Confounding
Assessing robustness to unmeasured confounding is critical for observational studies. Convex optimization enhances this assessment when multiple outcomes are involved. The method provides a systematic way to conduct sensitivity analysis. It evaluates how sensitive conclusions are to unmeasured confounders. This is done across a range of outcome variables. The approach allows researchers to understand the collective impact of unmeasured biases. It provides a more complete picture of study robustness. This is particularly important for complex interventions or exposures. These often affect several health outcomes. The technique helps to identify which findings are most vulnerable to confounding bias. It also highlights those that remain robust under various assumptions. This ensures a transparent and rigorous evaluation. It supports stronger claims regarding causal inference.
V.Sensitivity Analysis for Continuous Outcomes Treatment Effects
Sensitivity analysis extends to continuous outcome variables for assessing average treatment effects. This is a vital component of causal inference in observational studies. Continuous outcomes, such as blood pressure or income, are common. Evaluating the robustness of treatment effect estimates requires specialized methods. Optimization techniques provide solutions for these scenarios. They allow for a thorough examination of how unmeasured factors could alter conclusions. The analysis can proceed with or without assuming a known direction of effect for these confounders. This flexibility is crucial for real-world applications. It allows researchers to explore a wider range of hypothetical confounding scenarios. This systematic approach strengthens the validity of causal claims. It addresses potential biases inherent in observational data. This is a key aspect of modern bias reduction techniques.
5.1. Average Treatment Effect with Continuous Variables
Estimating the average treatment effect with continuous outcome variables is a primary goal. Sensitivity analysis is essential for these estimates in observational studies. It determines how susceptible observed effects are to unmeasured confounding. The methods explore a spectrum of confounding strengths. This helps quantify the potential impact of hidden biases. The analysis focuses on the change in the average treatment effect. This change occurs under various assumptions about unmeasured confounders. This provides a clear understanding of the study's limitations. It also highlights the range of plausible effects. This systematic approach contributes to more reliable causal inference. It helps researchers present a balanced view of their findings. This advanced analytical step is critical for drawing robust conclusions.
5.2. Incorporating Assumed Direction of Effect in Analysis
Sensitivity analysis for continuous outcomes can incorporate assumptions about the direction of effect. Sometimes, prior knowledge suggests the direction of an unmeasured confounder's influence. This information can be integrated into the analysis. For example, a confounder might be known to either inflate or deflate the observed treatment effect. Explicitly modeling this direction provides more focused sensitivity bounds. This refines the assessment of robustness. It allows for a more powerful and targeted examination of potential biases. When no such knowledge exists, analysis proceeds without directional assumptions. This ensures broad applicability. The ability to customize sensitivity analysis enhances its utility. It provides more precise insights into the reliability of causal inference. This flexibility is crucial for comprehensive bias reduction techniques.
Mục lục chi tiết luận án
Tải xuống file đầy đủ để xem toàn bộ nội dung
Tải đầy đủ (177 trang)Trích đoạn nội dung luận án
Tải xuống để đọc toàn bộUniversity of Pennsylvania ScholarlyCommons Publicly Accessible Penn Dissertations 2016 Modern Optimization in Observational Studies Colin Burton Fogarty University of Pennsylvania, colin.com Follow this and additional works at: https://repository.edu/edissertations Part of the Statistics and Probability Commons Recommended Citation Fogarty, Colin Burton, "Modern Optimization in Observational Studies" (2016). Publicly Accessible Penn Dissertations.edu/edissertations/1720 This paper is posted at ScholarlyCommons.edu/edissertations/1720 For more information, please contact repository@pobox. Modern Optimization in Observational Studies Abstract Perhaps the best known use of modern techniques for optimization in observational studies is within matching algorithms, wherein treated units are placed into matched sets with similar control units to adjust for overt biases. While the intuitive appeal of matching has been long understood, its ascent in popularity can be attributed in large part to computational advances in network flow optimization.
This dissertation explores how modern optimization can be leveraged to address other problems in observational studies. First, we demonstrate how, in the absence of covariate overlap, the maximal box problem can be used to define an interpretable study population wherein inference can be conducted without extrapolating on important variables. Next, we discuss how integer programming can be used to perform inference, construct confidence intervals, and provide sensitivity analyses for meaningful causal estimands in matched observational studies when the outcomes of interest are binary. Third, we present a method utilizing convex optimization for conducting a sensitivity analysis when there are multiple outcome variables of interest which, we show, can help attenuate the loss in power from accounting for multiple comparisons when assessing the robustness of a study's findings to unmeasured confounding.
Finally, we present methods for conducting a sensitivity analysis for the average treatment effect with continuous outcome variables with and without assuming a known direction of effect. Degree Type Dissertation Degree Name Doctor of Philosophy (PhD) Graduate Group Statistics First Advisor Dylan S. Small Keywords Causal Inference, Integer Programming, Matching, Observational Studies, Randomization Inference, Sensitivity Analysis Subject Categories Statistics and Probability This dissertation is available at ScholarlyCommons: https://repository.edu/edissertations/1720 MODERN OPTIMIZATION IN OBSERVATIONAL STUDIES Colin B. Fogarty A DISSERTATION in Statistics For the Graduate Group in Managerial Science and Applied Economics Presented to the Faculties of the University of Pennsylvania in Partial Fulfillment of the Requirements for the Degree of Doctor of Philosophy 2016 Supervisor of Dissertation Dylan S.
Small Professor of Statistics Graduate Group Chairperson Eric T. Chao Professor, Professor of Marketing, Statistics, and Education Dissertation Committee Paul R. Putzel Professor, Professor of Statistics Andreas Buja, Liem Sioe Liong/First Pacific Company Professor, Professor of Statistics MODERN OPTIMIZATION IN OBSERVATIONAL STUDIES c COPYRIGHT 2016 Colin Burton Fogarty This work is licensed under the Creative Commons Attribution NonCommercial-ShareAlike 3.0 License To view a copy of this license, visit http://creativecommons.org/licenses/by-nc-sa/3.0/ For Beatrice and Janice iii ACKNOWLEDGEMENT I would like to start by thanking my advisor, Dylan. I am simultaneously indebted to and in awe of the care and dedication given by you to each and every one of your students.
Time and time again I have been impressed by your seemingly boundless knowledge of the literature, your insights into the benefits and limitations of existing methods, and your unwavering love of scholarship. Your passion for research and advising is nothing short of inspirational, and I am proud to have worked with and learned from you over these past five years. I would next like to thank my committee members, Andreas and Paul. Andreas, the two classes that I took from you were fundamental in shaping my statistical intuition.
In his poem Maud Muller, John Greenleaf Whittier writes that “for of all sad words of tongue or pen, The saddest are these: ‘It might have been!’" While applicable in most walks of life, I now realize that hope springs eternal when these words are considered in the context of “dataset to dataset" variation and statistical inference. Paul, learning from your writings on observational studies has instilled within me the virtues of clarity, precision and conviction in writing. Each sentence should have a purpose, each theorem a necessity. I am also grateful for your kindness and willingness to meet with me to discuss and share ideas.
Given the contents of this dissertation, it goes without saying that your contributions to the field have had a profound impact on my thinking and interests. I would also like to thank the entirety of the Wharton Statistics Department for creating such a welcoming environment and making my five years as a PhD student so enjoyable. To the many professors with whom I have interacted - thank you for your time, your friendliness and for sharing your insights and perspectives. To our wonderful staff - in short, thank you for making my life so easy, be it through scheduling, reserving rooms, facilitating recom- mendation letters, help with computing, help with funding, or any of the other myriad ways you go above and beyond.
To my cohort, Ville, Kory, Tung, Julie, and Justin - thank you for your friendship, and thank you for your willingness to collaborate as we went through iv courses together. I have learned so much from each and every one of you. To the rest of the students with whom I have overlapped - thank you for your camaraderie, for your encouragement, and for your willingness to unwind after periods of hard work. Thanks and appreciation are, of course, also in order for my family.
Thank you so much for your love and encouragement throughout the years. Thank you for providing an environment which fostered independence while making it obvious that help was only a call away. Thank you for always being there for me through times of joy and times of hardship. No matter what life has thrown, and may throw, my way, I know I have and will always have your love and support.
Finally, to my loving wife, Beatrice. You are my inspiration and my motivation. You are the limitless source of positivity that drives me to be the best person I can be. Thank you for everything you do, and for everything you are.
v ABSTRACT MODERN OPTIMIZATION IN OBSERVATIONAL STUDIES Colin B. Small Perhaps the best known use of modern techniques for optimization in observational studies is within matching algorithms, wherein treated units are placed into matched sets with sim- ilar control units to adjust for overt biases. While the intuitive appeal of matching has been long understood, its ascent in popularity can be attributed in large part to computational advances in network flow optimization. This dissertation explores how modern optimization can be leveraged to address other problems in observational studies.
First, we demonstrate how, in the absence of covariate overlap, the maximal box problem can be used to define an interpretable study population wherein inference can be conducted without extrapolating on important variables. Next, we discuss how integer programming can be used to perform inference, construct confidence intervals, and provide sensitivity analyses for meaningful causal estimands in matched observational studies when the outcomes of interest are binary. Third, we present a method utilizing convex optimization for conducting a sensitivity analy- sis when there are multiple outcome variables of interest which, we show, can help attenuate the loss in power from accounting for multiple comparisons when assessing the robustness of a study’s findings to unmeasured confounding. Finally, we present methods for conducting a sensitivity analysis for the average treatment effect with continuous outcome variables with and without assuming a known direction of effect.
vi TABLE OF CONTENTS ACKNOWLEDGEMENT. vi LIST OF TABLES. x LIST OF ILLUSTRATIONS. xii CHAPTER 1 : Introduction.
1 CHAPTER 2 : Discrete Optimization for Interpretable Study Populations and Ran- domization Inference in an Observational Study of Severe Sepsis Mor- tality .2 Review of Causal Inference via Matching .3 Lack of Common Support .4 Defining a Study Population .5 Randomization Inference for the Average Treatment Effect with Binary Out- comes .6 Inference for Severe Sepsis Mortality. 30 CHAPTER 3 : Randomization Inference and Sensitivity Analysis for Composite Null Hypotheses with Binary Outcomes in Matched Observational Studies 33 3.2 Causal Inference after Matching .3 Composite Null Hypotheses .5 Inference and Sensitivity Analysis. 60 CHAPTER 4 : Sensitivity Analysis for Multiple Comparisons in Matched Observa- tional Studies through Quadratically Constrained Linear Programming 62 4.2 Notation for a Matched Observational Study .3 Sensitivity Analysis for Overall Significance .4 Improving Power through Quadratically Constrained Linear Programming .5 Familywise Error Control for Individual Null Hypotheses .6 Simulation Study: Gains in Power of a Sensitivity Analysis .7 Improved Robustness to Unmeasured Confounding for Elevated Napthalene in Smokers. 86 CHAPTER 5 : Sensitivity Analysis for the Average Treatment Effect in Matched Observational Studies .2 A Paired Observational Study .3 The Average Treatment Effect .4 Sensitivity Analysis for the Average Treatment Effect .5 Known Direction of Effect .6 Bigger Effect for Individuals More Likely to Receive Treatment .7 Simulation: The Impact of Assumptions on Sensitivity to Unmeasured Con- founding.
152 ix LIST OF TABLES TABLE 1 : Covariate Means and Standard Deviations, Original Population and Study Population for Tier 1 Covariates. 7 TABLE 2 : Estimated Differences in Severe Sepsis Mortality between ICU and Hospital Ward Patients in Study Population. 30 TABLE 3 : Computation Times for Testing Nulls on Risk Difference and Risk Ratio through Integer Programming. 55 TABLE 4 : The Impact of a Known Direction of Effect on Sensitivity Analyses.
58 TABLE 5 : Sensitivity Analysis for the Effect Ratio under Various Assumptions. 59 TABLE 6 : Power of a Sensitivity Analysis for the Overall Null. 80 TABLE 7 : Power of Closed Testing for Individual Nulls. 82 TABLE 8 : Worst-Case Confounders in a Particular Pair at Γ = 10 with Multiple Outcomes.
84 TABLE 9 : Means and Standard Deviations for Non-Binary Covariates Before Matching, Original Population and Study Population. 107 TABLE 10 : Percentages for Binary Covariates Before Matching, Original Popu- lation and Study Population. 108 TABLE 11 : Percentages of Missing Values, Original Population and Study Pop- ulation. 109 TABLE 12 : Computation Times for Testing Nulls on Risk Difference and Risk Ratio through Integer Programming using Acute Rehabilitation Data 136 TABLE 13 : Strong Familywise Error Control of Proposed Method through Closed Testing.
144 x LIST OF ILLUSTRATIONS FIGURE 1 : Lack of Common Support and the Maximal Box. 14 FIGURE 2 : Covariate Imbalances Before and After Full Matching, Study Pop- ulation. 24 FIGURE 3 : A Direct Acyclic Graph Illustrating a Sensitivity Analysis with Mul- tiple Outcomes. 64 FIGURE 4 : Power of a Sensitivity Analysis for the Average Treatment Effect.
102 FIGURE 5 : Proportion of Individuals Identified by the Method of King and Zeng (2006) as within the Area of Common Support. 110 FIGURE 6 : Randomization Distribution of the Average Treatment Effect at the Worst-Case Null Distribution. 115 FIGURE 7 : Standardized Differences Before and After Matching: Acute Reha- bilitation Study. 121 FIGURE 8 : Optimization Time as a Function of Matched Sets and Variables.
126 FIGURE 9 : Optimization Time as a Function of Matched Triples and Variables 127 FIGURE 10 : Optimization Time and the Degree of Allowed Unmeasured Con- founding. 128 FIGURE 11 : Optimization Time and the Null Hypothesis. 129 FIGURE 12 : The Impact of Overall Event Frequency on Optimization Time and the Number of Variables. 130 FIGURE 13 : The Impact of Event Frequency under Treatment on Optimization Time and the Number of Variables.
131 FIGURE 14 : The Impact of Event Frequency under Control on Optimization Time and the Number of Variables. 132 xi FIGURE 15 : Standardized Differences Before and After Matching: Smoking and Naphthalene Study. 142 xii CHAPTER 1 : Introduction In an ideal world there would be no need for observational studies; any hypothesized causal relationship would be tested through controlled randomized experiments, with randomiza- tion conferring both a “reasoned basis for inference” (Fisher, 1935) and protection against unmeasured confounding.
Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ
Trích dẫn luận án này
Colin Burton Fogarty (2016). Modern optimization in observational studies [Luận án tiến sĩ, University of Pennsylvania]. LuanAn.net. https://luanan.net/toan-hoc/toan-ung-dung/modern-optimization-in-observational-studies
Câu hỏi thường gặp
Luận án "Modern optimization in observational studies" nghiên cứu về vấn đề gì?
"Phân tích và tối ưu hóa phương pháp nghiên cứu quan sát bằng các kỹ thuật hiện đại để cải thiện độ chính xác và hiệu suất của kết quả."
Luận án "Modern optimization in observational studies" được bảo vệ tại trường nào?
Luận án này được bảo vệ tại University of Pennsylvania. Năm bảo vệ: 2016.
Luận án "Modern optimization in observational studies" thuộc chuyên ngành gì?
Luận án "Modern optimization in observational studies" thuộc chuyên ngành Statistics and Probability. Danh mục: Toán Ứng Dụng.
Luận án "Modern optimization in observational studies" có bao nhiêu trang?
Luận án "Modern optimization in observational studies" có 177 trang. Bạn có thể xem trước một phần tài liệu ngay trên trang web trước khi tải về.
Cách tải luận án "Modern optimization in observational studies" về máy như thế nào?
Để tải luận án về máy, bạn nhấn nút "Tải xuống ngay" trên trang này, sau đó hoàn tất thanh toán phí lưu trữ. File sẽ được tải xuống ngay sau khi thanh toán thành công. Hỗ trợ qua Zalo: 0559 297 239.