О калькуляторе: Linear regression calculator (line of best fit)
Least squares picks the line or curve that makes the sum of squared vertical distances from the points as small as possible. For a straight line the slope is b₁ = Sxy/Sxx and the intercept b₀ = ȳ − b₁x̄; polynomial fits of degree 2 to 4 solve the normal equations, and the exponential model fits a straight line to ln y. Each coefficient gets a standard error, a t statistic and a p-value, and R² gives the share of the variation in y that the fit accounts for.
The default relates weekly study hours to exam scores for 10 students. The fitted line is y = 48.89 + 3.109x, so each extra hour goes with about 3.1 more points, R² is 0.9933, and at 7.5 hours the line predicts 72.2.
The standard errors and p-values assume independent residuals with constant spread. A curved or funnel-shaped residual plot means the model is missing something, and predictions outside the range of the x data are flagged as extrapolations.
Вопросы
How do you calculate the line of best fit?
Find the means x̄ and ȳ, then the slope b₁ = Σ(x − x̄)(y − ȳ) / Σ(x − x̄)² and the intercept b₀ = ȳ − b₁x̄. For the default study-hours data, Sxy = 256.5 and Sxx = 82.5, so b₁ = 3.109 and b₀ = 48.89. This least-squares line minimizes the sum of squared vertical distances, as set out in the NIST/SEMATECH e-Handbook (§4.1.4.1).
What does R² tell you in regression?
R² is the fraction of the variation in y that the model accounts for: R² = 1 − SSE/SST. The default fit has R² = 0.9933, so study hours account for 99.3% of the spread in exam scores. For a straight line R² equals the square of the correlation r. A high R² does not show that x causes y, that the model has the right shape, or that predictions beyond the data will hold.
What is the difference between R² and adjusted R²?
Adjusted R² charges for each extra coefficient: 1 − (1 − R²)(n − 1)/(n − k), where k counts the coefficients including the intercept. R² never falls when a term is added, so a quartic always fits at least as well as a quadratic; adjusted R² falls when the new term explains less than it costs. For the default data, n = 10 and k = 2 turn R² = 0.99325 into 0.99241.
What does the p-value of the slope mean?
It tests the null hypothesis that the true slope is 0, using t = b₁/SE(b₁) with n − 2 degrees of freedom. For the default data t = 3.109/0.0906 = 34.3 with 8 df, giving p = 5.7 × 10⁻¹⁰: a slope this steep would be very unlikely if hours and scores had no linear relation. The p-value says nothing about how large or important the effect is.
When should I use an exponential fit instead of a straight line?
When y grows or shrinks by a roughly constant percentage for each unit of x, so that ln y against x is straight. The exponential example, with y rising from 3.0 to 22.1 as x goes from 0 to 5, fits y = 2.98·e^(0.4007x), an increase of 49.3% per unit of x. The fit is made on ln y, as Excel's LOGEST function does, so R² describes ln y rather than y.
Насколько точен «Linear regression calculator (line of best fit)»?
Точность зависит от введённых данных и допущений метода. Десятичная арифметика использует 50 значащих цифр, но оценки, численные методы и исходные данные могут быть менее точными; округление на экране не устраняет эти ограничения. Решённые примеры, проверенные по независимым источникам: 5. Например, «Exact straight line y = 1 + 2x» проверяется по источнику Constructed points (1,3), (2,5), (3,7) lie exactly on y = 1 + 2x; at x = 4, y = 9.
Откуда взята методика?
NIST/SEMATECH e-Handbook of Statistical Methods, §4.1.4.1 Linear least squares regression; NIST Statistical Reference Datasets — linear regression; Draper & Smith, Applied Regression Analysis, 3rd ed. (Wiley, 1998), ch. 1 and 12.