When:
Friday, October 25, 2024
11:00 AM - 12:00 PM CT
Where: Chambers Hall, Ruan Conference Room – lower level, 600 Foster St, Evanston, IL 60208 map it
Audience: Faculty/Staff - Student - Post Docs/Docs - Graduate Students
Cost: free
Contact:
Kisa Kowal
(847) 491-3974
Group: Department of Statistics and Data Science
Category: Academic, Lectures & Meetings
Subsampling for Big Data Regression with Measurement Constraints
Lin Wang, Assistant Professor of Statistics, Purdue University
Abstract: Despite the availability of extensive data sets, it is often impractical to observe the labels for all data points due to various measurement constraints in many applications. To address this challenge, subsampling approaches can be employed to select a subset of design points from a large pool for observation, resulting in substantial savings in labeling costs. In this presentation, I will introduce our recent research on computationally feasible subsampling techniques. Our primary focus is on regression with labeled data, which includes linear regression, ridge regression, and nonparametric additive regression. For these regression tasks, we have developed sampling approaches that aim to minimize the mean squared error in estimations and predictions. We will demonstrate the effectiveness of our proposed approaches through theoretical analysis and extensive numerical results.