Determination of the Representative Sample Size in Linear Regression
Anatoly Rayev
Abstract
Very often, accuracy of analysis and forecasting (multiple coefficient of regression and residual means) obtained for a sample used to formulate a regression model is not equal to the accuracy achieved for another homogeneous sample. Indeed, accuracy of analysis and forecasting based on another sample is much worse. This is explained by the discrepancy between a postulated and a real model. This discrepancy is caused by the redundancy of the number of terms of an approximating series. It describes a phenomenon under study with the sample noise. To filter this noise it is necessary to correctly choose the terms of an approximating series and to determine their number, which in turn depends on sample size. It is possible to include in this approximating series not only independent variables but also various functions of the value of the independent variables (e.g., squares, cubes, algorithms, etc.). This paper gives the procedure for selection of an approximating series.
Create a lesson
Related papers
Minimax optimality for sequential gradient-free minimization of smooth functions and their derivatives
Théo Paquier, Alexandre B Tsybakov, François Portier et al.
Randomization Inference with Concentration Inequalities
Tobias Freidling
On the continuity of the Tukey depth function for fuzzy data
Luis González-De La Fuente, Alicia Nieto-Reyes, Pedro Terán
Recursive-Head Geometry and Order-Free Efficient Inference in Finite-State Nested Markov Models
Haoyu Wei
Finite-Sample Hausdorff Bounds and Hadamard Sensitivity for Regressions with MNAR Covariates
Hugo Dunias
Semiparametric Efficient Inference under Non-Informative Complex Survey Designs
Hiroki Chiba, Kosuke Morikawa