Gary King: Giải pháp suy luận sinh thái - Tái tạo hành vi cá nhân từ dữ liệu tổng hợp

Phần 1: Giới thiệu tổng quan về chủ đề. Khám phá các khái niệm cốt lõi, tầm quan trọng và mục tiêu nghiên cứu.

Chuyên ngành

Political Science - Statistical Methods

Tác giả

Luan An

Thể loại

Sách

Năm xuất bản

Số trang

54

Thời gian đọc

9 phút

Lượt xem

0

Lượt tải

0

Phí lưu trữ

40 Point

Tổng quan nhanh

Chủ đề:
Tổng quan Khái niệm Vấn đề Suy luận Sinh thái
Số trang:
54 trang
Chuyên ngành:
Political Science - Statistical Methods
Tác giả:
Năm:

Tóm tắt nội dung luận án

I.Tổng quan Khái niệm Vấn đề Suy luận Sinh thái

Đây là phần giới thiệu. Nó cung cấp một cái nhìn tổng quan. Vấn đề suy luận sinh thái được trình bày. Mục tiêu là tái tạo hành vi cá nhân. Việc này dựa trên dữ liệu tổng hợp. Đây là thách thức phức tạp. Cần hiểu rõ các khái niệm ban đầu. Phần mở đầu này đặt nền tảng. Giúp độc giả nắm bắt vấn đề. Thuật ngữ cơ bản được giải thích. Việc bắt đầu giải quyết vấn đề đòi hỏi sự chuẩn bị kỹ lưỡng. Tổng quan này là bước đầu tiên quan trọng.

1.1. Khái niệm ban đầu Từ dữ liệu tổng hợp

Vấn đề suy luận sinh thái bắt đầu. Nó liên quan đến việc hiểu hành vi. Cụ thể là hành vi ở cấp độ cá nhân. Tuy nhiên, dữ liệu có sẵn chỉ ở dạng tổng hợp. Ví dụ, tỷ lệ bỏ phiếu của một khu vực. Không phải dữ liệu bỏ phiếu của từng người. Thách thức là làm sao suy luận ngược. Từ tổng thể đến cá nhân. Đây là một nguyên tắc cơ bản cần nắm vững.

1.2. Phần mở đầu Nhu cầu thực tiễn

Nhu cầu giải quyết vấn đề này rất lớn. Nó xuất hiện trong nhiều lĩnh vực nghiên cứu. Khoa học chính trị thường xuyên gặp phải. Xã hội học và y tế cũng tương tự. Dữ liệu cá nhân thường khó tiếp cận. Hoặc vì lý do riêng tư. Hoặc do chi phí thu thập cao. Do đó, việc sử dụng dữ liệu tổng hợp là bắt buộc. Suy luận sinh thái cung cấp giải pháp cần thiết.

II.Giới thiệu Tổng quát về Suy luận Sinh thái Cơ bản

Phần này giới thiệu tổng quát hơn. Nó đi sâu vào bản chất cơ bản. Của vấn đề suy luận sinh thái. Mục tiêu chính là tái cấu trúc hành vi. Từ thông tin nhóm lớn. Quá trình này phức tạp. Nó đòi hỏi phương pháp luận chặt chẽ. Đảm bảo độ chính xác cao. Chuẩn bị cẩn thận là yếu tố then chốt. Tổng quan này cung cấp cái nhìn ban đầu. Về cách tiếp cận vấn đề cốt lõi. Đây là các bước đầu tiên trong giải pháp.

2.1. Các bước đầu tiên Tái cấu trúc hành vi

Mục tiêu chính luôn là hành vi cá nhân. Các nhà nghiên cứu muốn hiểu. Cách mỗi người đưa ra quyết định. Tuy nhiên, dữ liệu chỉ cho thấy xu hướng nhóm. Tái cấu trúc là quá trình gian nan. Nó đòi hỏi công cụ thống kê mạnh mẽ. Để biến dữ liệu nhóm thành thông tin cá nhân. Đây là một trong những nguyên tắc cơ bản. Giúp bắt đầu phân tích hiệu quả.

2.2. Nguyên tắc cơ bản Vượt qua giới hạn dữ liệu

Dữ liệu tổng hợp có giới hạn riêng. Nó không thể tiết lộ tất cả. Thông tin về từng cá nhân. Nguyên tắc là phát triển mô hình. Mô hình này phải có khả năng. Giải thích mối quan hệ tiềm ẩn. Giữa dữ liệu ở cấp độ nhóm. Và hành vi ở cấp độ cá nhân. Đây là một khái niệm ban đầu quan trọng. Giúp vượt qua rào cản thông tin.

III.Tầm quan trọng của Suy luận Sinh thái Các Bước Đầu Tiên

Suy luận sinh thái có tầm quan trọng lớn. Nó giải quyết thiếu hụt dữ liệu. Đặc biệt khi dữ liệu cá nhân hiếm. Hoặc không thể tiếp cận được. Các bước đầu tiên đã được đề xuất. Cung cấp một phương pháp đáng tin cậy. Để phân tích dữ liệu tổng hợp. Nhằm đạt được cái nhìn sâu sắc. Về hành vi cá nhân. Đây là phần mở đầu cho giải pháp. Giúp các nhà nghiên cứu chuẩn bị tốt hơn. Cho việc áp dụng thực tế.

3.1. Nhu cầu cấp thiết Dữ liệu thiếu hụt

Sự thiếu hụt dữ liệu cá nhân là phổ biến. Nhiều chính phủ và tổ chức. Chỉ công bố dữ liệu tổng hợp. Dữ liệu này có sẵn rộng rãi hơn. Nhưng nó đặt ra thách thức. Làm thế nào để khai thác giá trị? Suy luận sinh thái lấp đầy khoảng trống này. Nó cung cấp công cụ cần thiết. Cho việc nghiên cứu chính xác. Đây là một trong những nguyên tắc cơ bản.

3.2. Phương pháp tiếp cận ban đầu Giải pháp Gary King

Gary King đã đề xuất giải pháp đột phá. Nó giải quyết triệt để vấn đề. Của suy luận sinh thái. Phương pháp này tập trung vào tái tạo. Hành vi cá nhân một cách có hệ thống. Dựa trên dữ liệu tổng hợp hiện có. Giải pháp này mở ra hướng mới. Cho nghiên cứu thực nghiệm. Đây là một khái niệm ban đầu quan trọng. Để bắt đầu hiểu cách giải quyết vấn đề.

IV.Thiết lập Ban đầu Vấn đề Phân tích Dữ liệu Tổng hợp

Phần này thiết lập vấn đề chính thức. Nó định nghĩa rõ ràng thách thức. Của việc phân tích dữ liệu tổng hợp. Để suy luận về cá nhân. Đây là một bước chuẩn bị quan trọng. Đảm bảo hiểu đúng phạm vi. Và các giới hạn của vấn đề. Phân tích này là nền tảng. Cho các phương pháp giải quyết sau. Các nguyên tắc cơ bản được nhấn mạnh. Giúp các nhà nghiên cứu bắt đầu hiệu quả.

4.1. Định nghĩa chính thức Vấn đề cốt lõi

Vấn đề suy luận sinh thái được định nghĩa. Nó là việc ước tính các đại lượng. Các đại lượng này thuộc cấp độ cá nhân. Nhưng chỉ từ dữ liệu cấp độ nhóm. Đây là một bài toán thống kê phức tạp. Nó đòi hỏi mô hình hóa cẩn thận. Và kiểm tra nghiêm ngặt. Định nghĩa này cung cấp một khái niệm ban đầu rõ ràng. Về bản chất của thách thức.

4.2. Các yếu tố chuẩn bị Dữ liệu đầu vào

Dữ liệu đầu vào là trọng tâm. Nó luôn là dữ liệu tổng hợp. Ví dụ: kết quả bầu cử theo khu vực. Mục tiêu là suy luận về bỏ phiếu cá nhân. Quá trình chuẩn bị dữ liệu rất quan trọng. Cần xác định các biến số liên quan. Và đảm bảo chất lượng dữ liệu. Đây là các bước đầu tiên thiết yếu. Để thành công trong suy luận sinh thái.

Xem trước tài liệu
Tải đầy đủ để xem toàn bộ nội dung
Part1

Tải xuống file đầy đủ để xem toàn bộ nội dung

Tải đầy đủ (54 trang)

Trích đoạn nội dung luận án

Tải xuống để đọc toàn bộ

Gary King: A Solution to the Ecological Inference Problem: Reconstructing Individual Behavior from Aggregate Data is published by Princeton University Press and copyrighted,  1997, Princeton University Press. All rights reserved. This text may be used and shared in accordance with the fair-use provisions of US copyright law, and it may be archived and redistributed in electronic form, provided that this notice is carried, Princeton University Press is notified, the entire original is distributed without modification, and no fee is charged for access. Archiving, redistribution, or republication of this text on other terms, in any medium, requires the consent of Princeton University Press.

For COURSE PACK PERMISSIONS, refer to entry on previous menu. For more information, send e-mail to permissions@pupress.edu A Solution to the Ecological Inference Problem A Solution to the Ecological Inference Problem reconstructing individual behavior from aggregate data Gary King PRINCETON UNIVERSITY PRESS P R I N C E T O N, N E W J E R S E Y Copyright © 1997 by Princeton University Press Published by Princeton University Press, 41 William Street, Princeton, New Jersey 08540 In the United Kingdom: Princeton University Press, Chichester, West Sussex All Rights Reserved Library of Congress Cataloging-in-Publication Data King, Gary. A solution to the ecological inference problem: reconstructing individual behavior from aggregate data / Gary King. Includes bibliographical references and index.

Political science—Statistical methods.072—dc20 9632986 CIP This book has been composed in Palatino Princeton University Press books are printed on acid-free paper and meet the guidelines for permanence and durability of the Committee on Production Guidelines for Book Longevity of the Council on Library Resources Printed in the United States of America by Princeton Academic Press 1 3 5 7 9 10 8 6 4 2 1 3 5 7 9 10 8 6 4 2 (Pbk.) For Ella Michelle King Contents List of Figures xi List of Tables xiii Preface xv Part I: Introduction 1 1 Qualitative Overview 3 1.1 The Necessity of Ecological Inferences 7 1.5 The Method 26 2 Formal Statement of the Problem 28 Part II: Catalog of Problems to Fix 35 3 Aggregation Problems 37 3.1 Goodman’s Regression: A Definition 37 3.2 The Indeterminacy Problem 39 3.3 The Grouping Problem 46 3.4 Equivalence of the Grouping and Indeterminacy Problems 53 3.5 A Concluding Definition 54 4 Non-Aggregation Problems 56 4.1 Goodman Regression Model Problems 56 4.2 Applying Goodman’s Regression in 2 × 3 Tables 68 4.3 Double Regression Problems 71 4.4 Concluding Remarks 73 Part III: The Proposed Solution 75 5 The Data: Generalizing the Method of Bounds 77 5.1 Homogeneous Precincts: No Uncertainty 78 viii Contents 5.2 Heterogeneous Precincts: Upper and Lower Bounds 79 5.1 Precinct-Level Quantities of Interest 79 5.2 District-Level Quantities of Interest 83 5.3 An Easy Visual Method for Computing Bounds 85 6 The Model 91 6.1 The Basic Model 92 6.1 Observable Implications of Model Parameters 96 6.2 Parameterizing the Truncated Bivariate Normal 102 6.3 Computing 2p Parameters from Only p Observations 106 6.4 Connections to the Statistics of Medical and Seismic Imaging 112 6.5 Would a Model of Individual-Level Choices Help? 119 7 Preliminary Estimation 123 7.2 The Likelihood Function 132 7.5 Summarizing Information about Estimated Parameters 139 8 Calculating Quantities of Interest 141 8.1 Simulation Is Easier than Analytical Derivation 141 8.1 Definitions and Examples 142 8.2 Simulation for Ecological Inference 144 8.2 Precinct-Level Quantities 145 8.3 District-Level Quantities 149 8.4 Quantities of Interest from Larger Tables 151 8.1 A Multiple Imputation Approach 151 8.2 An Approach Related to Double Regression 153 8.5 Other Quantities of Interest 156 9 Model Extensions 158 9.1 What Can Go Wrong? 158 9.2 Incorrect Distributional Assumptions 161 9.2 Avoiding Aggregation Bias 168 9.1 Using External Information 169 Contents ix 9.2 Unconditional Estimation: Xi as a Covariate 174 9.3 Tradeoffs and Priors for the Extended Model 179 9.4 Ex Post Diagnostics 183 9.3 Avoiding Distributional Problems 184 9.2 A Nonparametric Approach 191 Part IV: Verification 197 10 A Typical Application Described in Detail: Voter Registration by Race 199 10.3 Computing Quantities of Interest 207 10.3 Other Quantities of Interest 215 11 Robustness to Aggregation Bias: Poverty Status by Sex 217 11.1 Data and Notation 217 11.2 Verifying the Existence of Aggregation Bias 218 11.3 Fitting the Data 220 11.4 Empirical Results 222 12 Estimation without Information: Black Registration in Kentucky 226 12.3 Fitting the Data 228 12.4 Empirical Results 232 13 Classic Ecological Inferences 235 13.2 Black Literacy in 1910 241 Part V: Generalizations and Concluding Suggestions 247 14 Non-Ecological Aggregation Problems 249 14.1 The Geographer’s Modifiable Areal Unit Problem 249 x Contents 14.1 The Problem with the Problem 250 14.2 Ecological Inference as a Solution to the Modifiable Areal Unit Problem 252 14.2 The Statistical Problem of Combining Survey and Aggregate Data 255 14.3 The Econometric Problem of Aggregating Continuous Variables 258 14.4 Concluding Remarks on Related Aggregation Research 262 15 Ecological Inference in Larger Tables 263 15.1 An Intuitive Approach 264 15.2 Notation for a General Approach 267 15.4 The Statistical Model 271 15.6 Calculating the Quantities of Interest 276 15.7 Concluding Suggestions 276 16 A Concluding Checklist 277 Part VI: Appendices 293 A Proof That All Discrepancies Are Equivalent 295 B Parameter Bounds 301 B.2 Heterogeneous Precincts: β’s and θ’s 302 B.3 Heterogeneous Precincts: λi ’s 303 C Conditional Posterior Distribution 304 C.1 Using Bayes Theorem 305 C.2 Using Properties of Normal Distributions 306 D The Likelihood Function 307 E The Details of Nonparametric Estimation 309 F Computational Issues 311 Glossary of Symbols 313 References 317 Index 337 Figures 1.1 Model Verification: Voter Turnout among African Americans in Louisiana Precincts 23 1.2 Non-Minority Turnout in New Jersey Cities and Towns 25 3.1 How a Correlation between the Parameters and Xi Induces Bias 41 4.1 Scatter Plot of Precincts in Marion County, Indiana: Voter Turnout for the U. Senate by Fraction Black, 1990 60 4.2 Evaluating Population-Based Weights 64 4.3 Typically Massive Heteroskedasticity in Voting Data 66 5.1 A Data Summary Convenient for Statistical Modeling 81 5.2 Image Plots of Upper and Lower Bounds on βbi 86 5.3 Image Plots of Upper and Lower Bounds on βw i 87 5.4 Image Plots of Width of Bounds 88 5.5 A Scattercross Graph of Voter Turnout by Fraction Hispanic 89 6.1 Features of the Data Generated by Each Parameter 100 6.2 Truncated Bivariate Normal Distributions 105 6.4 Truncated Bivariate Normal Surface Plot 116 7.1 Verifying Individual-Level Distributional Assumptions with Aggregate Data 126 7.2 Observable Implications for Sample Parameter Values 127 7.3 Likelihood Contour Plots 137 8.1 Posterior Distributions of Precinct Parameters βbi 148 8.2 Support of the Joint Distribution of θib and βbi with Bounds Specified for Drawing λbi 155 9.1 The Worst of Aggregation Bias: Same Truth, Different Observable Implications 160 9.2 The Worst of Distributional Violations: Different True Parameters, Same Observable Implications 163 9.3 Conclusive Evidence of Aggregation Bias from Aggregate Data 176 9.5 Controlling for Aggregation Bias 179 9.6 Extended Model Tradeoffs 180 9.7 A Tomography Plot with Evidence of Multiple Modes 187 9.8 Building a Nonparametric Density Estimate 194 9.9 Nonparametric Density Estimate for a Difficult Case 195 xii Figures 10.1 A Scattercross Graph for Southern Counties, 1968 201 10.2 Tomography Plot of Southern Race Data with Maximum Likelihood Contours 204 10.3 Scatter Plot with Maximum Likelihood Results Superimposed 206 10.4 Posterior Distribution of the Aggregate Quantities of Interest 208 10.5 Comparing Estimates to the Truth at the County Level 210 10.7 Verifying Uncertainty Estimates 213 10.8 275 Lines Fit to 275 Points 214 11.1 South Carolina Tomography Plot 221 11.2 Posterior Distributions of the State-Wide Fraction in Poverty by Sex in South Carolina 222 11.3 Fractions in Poverty for 3,187 South Carolina Block Groups 223 11.4 Percentiles at Which True Values Fall 224 12.1 A Scattercross Graph of Fraction Black by Fraction Registered 227 12.2 Tomography Plot with Parametric Contours and a Nonparametric Surface Plot 229 12.3 Posterior Distributions of the State-Wide Fraction of Blacks and Whites Registered 231 12.4 Fractions Registered at the County Level 232 12.5 80% Posterior Confidence Intervals by True Values 233 13.1 Fulton County Voter Transitions 236 13.2 Aggregation Bias in Fulton County Data 238 13.3 Fulton County Tomography Plot 239 13.4 Comparing Voter Transition Rate Estimates with the Truth in Fulton County 241 13.5 Alternative Fits to Literacy by Race Data 242 13.6 Black Literacy Tomography Plot and True Points 243 13.7 Comparing Estimates to the County-Level Truth in Literacy by Race Data 244 Tables 1.1 The Ecological Inference Problem at the District Level 13 1.2 The Ecological Inference Problem at the Precinct Level 14 1.3 Sample Ecological Inferences 16 2.1 Basic Notation for Precinct i 29 2.2 Alternative Notation for Precinct i 31 2.3 Simplified Notation for Precinct i 31 4.1 Comparing Goodman Model Parameters to the Parameters of Interest in the 2 × 3 Table 70 9.1 Consequences of Spatial Autocorrelation: Monte Carlo Evidence 168 9.2 Consequences of Distributional Misspecification: Monte Carlo Evidence 189 10.1 Maximum Likelihood Estimates 202 10.2 Reparameterized Maximum Likelihood Estimates 203 10.3 Verifying Estimates of ψ 207 11.1 Evidence of Aggregation Bias in South Carolina 219 11.2 Goodman Model Estimates: Poverty by Sex 220 12.1 Evidence of Aggregation Bias in Kentucky 228 12.2 80% Confidence Intervals for ψ̆ and ψ 230 15.1 Example of a Larger Table 265 15.2 Notation for a Large Table 268 Preface In this book, I present a solution to the ecological inference problem: a method of inferring individual behavior from aggregate data that works in practice. Ecological inference is the process of using aggre- gate (i., “ecological”) data to infer discrete individual-level relation- ships of interest when individual-level data are not available. Existing methods of ecological inference generate very inaccurate conclusions about the empirical world—which thus gives rise to the ecological in- ference problem.

Most scholars who analyze aggregate data routinely encounter some form of the this problem. The ecological inference problem has been among the longest standing, hitherto unsolved problems in quantitative social science. It was originally raised over seventy-five years ago as the first statistical problem in the nascent discipline of political science, and it has held back research agendas in most of its empirical subfields. Ecological inferences are required in political science research when individual- level surveys are unavailable (for example, local or comparative electoral politics), unreliable (racial politics), insufficient (political ge- ography), or infeasible (political history).

They are also required in numerous areas of major significance in public policy (for example, for applying the Voting Rights Act) and other academic disciplines, ranging from epidemiology and marketing to sociology and quanti- tative history.1 Because the ecological inference problem is caused by the lack of individual-level information, no method of ecological inference, including that introduced in this book, will produce precisely ac- curate results in every instance. However, potential difficulties are minimized here by models that include more available information, diagnostics to evaluate when assumptions need to be modified, and realistic uncertainty estimates for all quantities of interest. For po- litical methodologists, many opportunities remain, and I hope the 1 What is “ecological” about the aggregate data from which individual behavior is to be inferred? The name has been used at least since the late 1800s and stems from the word ecology, the science of the interrelationship of living things and their environ- ments. Statistical measures taken at the level of the environment, such as summaries of geographic areas or other aggregate units, are widely known as ecological data.

Eco- logical inference is the process of using ecological data to learn about the behavior of individuals within these aggregates. xvi Preface results reported here lead to continued research into and further improvements in the methods of ecological inference. But most im- portantly, the solution to the ecological inference problem presented here is designed so that empirical researchers can investigate sub- stantive questions that have heretofore proved intractable. Perhaps it will also lead to new theories and empirical research in areas where analysts have feared to tread due to the lack of reliable ecological methods or individual-level data.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Trích dẫn luận án này

Gary King (1997). Giải quyết vấn đề suy luận sinh thái: Hành vi cá nhân từ dữ liệu tổng hợp [Luận án tiến sĩ]. LuanAn.net. https://luanan.net/ly-luan-va-lich-su-giao-duc/triet-hoc-giao-duc/part1

Câu hỏi thường gặp

Luận án "Giải quyết vấn đề suy luận sinh thái: Hành vi cá nhân từ dữ liệu tổng hợp" nghiên cứu về vấn đề gì?

Phần 1: Giới thiệu tổng quan về chủ đề. Khám phá các khái niệm cốt lõi, tầm quan trọng và mục tiêu nghiên cứu.

Luận án "Giải quyết vấn đề suy luận sinh thái: Hành vi cá nhân từ dữ liệu tổng hợp" thuộc chuyên ngành gì?

Luận án "Giải quyết vấn đề suy luận sinh thái: Hành vi cá nhân từ dữ liệu tổng hợp" thuộc chuyên ngành Political Science - Statistical Methods. Danh mục: Triết Học Giáo Dục.

Luận án "Giải quyết vấn đề suy luận sinh thái: Hành vi cá nhân từ dữ liệu tổng hợp" có bao nhiêu trang?

Luận án "Giải quyết vấn đề suy luận sinh thái: Hành vi cá nhân từ dữ liệu tổng hợp" có 54 trang. Bạn có thể xem trước một phần tài liệu ngay trên trang web trước khi tải về.

Cách tải luận án "Giải quyết vấn đề suy luận sinh thái: Hành vi cá nhân từ dữ liệu tổng hợp" về máy như thế nào?

Để tải luận án về máy, bạn nhấn nút "Tải xuống ngay" trên trang này, sau đó hoàn tất thanh toán phí lưu trữ. File sẽ được tải xuống ngay sau khi thanh toán thành công. Hỗ trợ qua Zalo: 0559 297 239.

Luận án liên quan

Chia sẻ tài liệu: Facebook Twitter