Linear model
daspi.anova.model.LinearModel(source, target, factors, covariates=None, alpha=0.05, order=1, skip_intercept_as_least=True, generalized_vif=True, encode_factors=True, fit_at_init=True)
¶
Bases: BaseHTMLReprModel
This class is used to create and simplify linear models so that only significant factors describe the model.
Balanced models (DOEs or EVOPs) including continuous can be analyzed. With this class, you can create an encoded design matrix with all factor levels, including their interactions. All non-significant factors can then be automatically eliminated. Furthermore, this class allows the examination of main effects, the sum of squares (explained variation), and the ANOVA table in more detail.
| PARAMETER | DESCRIPTION |
|---|---|
source
|
Pandas DataFrame as tabular data in a long format used for the model.
TYPE:
|
target
|
Column name of the endogenous variable.
TYPE:
|
factors
|
Column names of the exogenous variables that can be actively changed (factor levels for DOE or EVOP). Interactions are also created for these variables if the order is set > 1.
TYPE:
|
covariates
|
Column names for exogenous variables that are logged but cannot be influenced. No interactions are created for these variables. Default is None, which means no covariates are used.
TYPE:
|
alpha
|
Threshold as alpha risk. All factors, including continuous and intercept, that have a p-value smaller than alpha are removed during the automatic elimination of the factors. Default is 0.05.
TYPE:
|
skip_intercept_as_least
|
If True, the intercept is not removed as a least significant
term when using recursive factor elimination. Also if True, the
intercept does not appear when calling the
TYPE:
|
encode_factors
|
If True, all of the provided factor variables are encoded using one-hot encoding by changing the data type to category. Otherwise they are interpreted as continuous variables when possible, by default True.
TYPE:
|
fit_at_init
|
If True, the model is fitted at initialization. Otherwise, the
model is fitted after calling the
TYPE:
|
Examples:
Do some ANOVA and statistics on a dataset. Run the example below in a Jupyther Notebook to see the results.
import daspi as dsp
df = dsp.load_dataset('painkillers-dissolution')
model = dsp.LinearModel(
source=df,
target='dissolution',
factors=['employee', 'stirrer', 'brand', 'catalyst', 'water'],
covariates=['temperature', 'preparation'],
order=2,
covariates=None)
# Store goodnes of fit values for each elimination step
df_gof = pd.concat(model.recursive_elimination())
# Plot residual and parameter relevance analysis
dsp.ResidualsCharts(model).plot().stripes().label(info=True)
dsp.ParameterRelevanceCharts(model).plot().label(info=True)
# Get HTML output
model
Notes
Always be careful when removing the intercept. If the intercept is missing, Patsy will automatically add all one-hot encoded levels for a categorical variable to compensate for this missing term. This will result in extremely high VIFs. That means that the parameters are not linearly independent.
Terminology
- factor: The original name as it appears in the provided data.
- term: The name of the term as it appears in the design matrix. All factor names are converted to "xi" where i is an ascending number. The name of the first factor becomes "x0", the second becomes "x1",... The covariate variables are also converted in the same way, but instead of the letter "x" the letter "e" is used.
- parameter: The variable names in connection with a coefficient, but in the composition with the original factor names.
covariates = covariates or []
instance-attribute
¶
The list of covariates variables used in the linear model.
target = target
instance-attribute
¶
The name of the target variable for the linear model.
factors = factors
instance-attribute
¶
The list of factors used in the linear model.
target_map = {target: 'y'}
instance-attribute
¶
A dictionary that maps the original factor names to the encoded factor names used in the model.
feature_map = {f: _f for f, _f in zip(self.factors, f_main_terms)} | {c: _c for c, _c in zip(self.covariates, d_main_terms)}
instance-attribute
¶
A dictionary that maps the original factor names to the encoded names (term) used in the model.
main_term_map = {v: k for k, v in self.feature_map.items()}
instance-attribute
¶
A dictionary that maps the main term names (no interactions) back to the original factor names used in the model.
skip_intercept_as_least = skip_intercept_as_least
instance-attribute
¶
If True, the intercept is not treated as a least significant term
generalized_vif = generalized_vif
instance-attribute
¶
If True, the generalized VIF is calculated when possible. Otherwise, only the straightforward VIF (via R2) is calculated.
excluded = set()
instance-attribute
¶
A set of factor names that should be excluded from the model.
data = source.rename(columns=self.feature_map | self.target_map)[list(self.feature_map.values()) + list(self.target_map.values())].copy()
instance-attribute
¶
The Pandas DataFrame containing the data for the linear model.
fitted
property
¶
Whether the model is fitted (read-only).
model
property
¶
Get regression results of fitted model. Calls fit method
if no model is fitted yet (read-only).
initial_formula
property
¶
Get the initial formula for the ANOVA model (read-only).
The initial formula is constructed from the target variable and the initial terms, which are the terms with an interaction order less than or equal to the specified order.
uncertainty
property
¶
Get uncertainty of the model as square root of MS_Residual (read-only).
alpha
property
writable
¶
Alpha risk as significance threshold for p-value of exegenous factors.
design_info
property
¶
Get the DesignInfo instance of current fitted model (read-only).
terms
property
¶
Get the encoded names of all variables for the current fitted model (read-only).
parameters
property
¶
Get the names of all variables for the current fitted model in the composition using the original factor and covariates names (read-only).
term_map
property
¶
Get the names of all internal used and original terms variables for the current fitted model as dict (read-only)
main_parameters
property
¶
Get all main parameters of current model excluding intercept (read-only).
formula
property
¶
Get the formula used for the linear model, excluding any
factors specified in the exclude attribute. The formula is
constructed based on the excluded factors. If the intercept is
excluded, the formula will include '-1' as the first term.
Otherwise, the formula will include the original terms
(excluding the excluded factors) separated by '+' (read-only).
effect_threshold
property
¶
Calculates the threshold for the effect of adding a term to the model. The threshold is calculated as the inverse survival function (inverse of sf) at the given alpha (read-only).
design_matrix
property
¶
Get the design matrix of the current fitted model (read-only).
is_hierarchical()
¶
Check if current fitted model is hierarchical.
effects(include_intercept=True)
¶
Calculates the impact of each term on the target. The effects are described as absolute number of the parameter coefficients devided by its standard error.
| PARAMETER | DESCRIPTION |
|---|---|
include_intercept
|
If False, the intercept is dropped from the returned Series, by default True.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
Series
|
A Pandas Series containing the effects of each term on the target variable. The index of the Series corresponds to the term names, and the values represent the calculated effects. |
p_values(include_intercept=True)
¶
Get P-value for significance of adding model terms using anova typ III table for current model.
| PARAMETER | DESCRIPTION |
|---|---|
include_intercept
|
If False, the intercept is dropped from the returned Series, by default True.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
Series
|
A Pandas Series containing the p-values of each term on the target variable. The index of the Series corresponds to the term names, and the values represent the calculated p-values. |
least_parameter()
¶
Get the parameter name with the least effect or the least p-value coming from a ANOVA typ III table of current fitted model.
| RETURNS | DESCRIPTION |
|---|---|
str
|
The parameter name with the least p_value if possible or the parameter name with the least effect. |
Notes
This method checks if any p-values are missing (NaN). If there are missing p-values, it returns the term name that has the smallest effect on the target variable. Otherwise, it returns the term name with the least p-value for the F-stats coming from current ANOVA table.
fit(**kwds)
¶
Create and fit a ordinary least squares model using current
formula. To fit with a user-defined formula, use the
formula keyword argument.
| PARAMETER | DESCRIPTION |
|---|---|
**kwds
|
Pass formula and other keyword arguments to
DEFAULT:
|
r2_pred()
¶
Calculate the predicted R-squared (R2_pred) for the fitted model.
| RETURNS | DESCRIPTION |
|---|---|
float
|
The predicted R-squared value for the fitted model. |
Notes
The predicted R-squared is a measure of how well the model would predict new observations. It is calculated as:
R2_pred = 1 - (Sum of Squared Prediction Errors / Total Sum of Squares)
Where the prediction errors are calculated as the residuals divided by (1 - leverage), where leverage is the diagonal elements of the projection matrix P = X(X'X)^(-1)X'.
References
Calculations are made according to: https://support.minitab.com/de-de/minitab/help-and-how-to/statistical-modeling/doe/how-to/factorial/analyze-factorial-design/methods-and-formulas/goodness-of-fit-statistics/#press
Further information about projection matrix (influence matrix or hat matrix) can be found in: https://en.wikipedia.org/wiki/Projection_matrix
gof_metrics(index=0)
¶
Get different goodness-of-fit metrics (read-only).
| PARAMETER | DESCRIPTION |
|---|---|
index
|
Value is set as index. When using the method recursive_elimination, the current step is passed as index
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
The goodness-of-fit metrics table as DataFrame containing the following columns: - 'formula' = current formula - 's' = Uncertainty of the model as square root of MS_Residual - 'aic' = Akaike's information criteria - 'r2' = R-squared of the model - 'r2_adj' = adjusted R-squared - 'least_parameter' = the least significant term - 'p_least' = The p-value of least significant term, coming from ANOVA table Type-III. - 'hierarchical' = True if model is hierarchical |
summary(anova_typ=None, vif=True, **kwds)
¶
Generate a summary of the fitted model.
| PARAMETER | DESCRIPTION |
|---|---|
anova_typ
|
If not None, add an ANOVA table of provided type to the summary, by default None.
TYPE:
|
vif
|
If True, variance inflation factors (VIF) are added to the anova table. Will only be considered if anova_typ is not None, by default True
TYPE:
|
**kwds
|
Additional keyword arguments to be passed to the
DEFAULT:
|
| RETURNS | DESCRIPTION |
|---|---|
Summary
|
A summary object containing information about the fitted model. |
eliminate(parameter)
¶
Removes the given parameter from the model by adding it to t
he exclude set. Call fit to refit the model.
| PARAMETER | DESCRIPTION |
|---|---|
parameter
|
The factor name, the covariates name or the interaction of multiple factors to be removed from the model.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
Self
|
The current instance of the model for more method chaining. |
Examples:
Prepare a LinearModel instance, fit the model and plot the relevance of the parameters. If you run the following code in a Jupyter Notebook, the plot and the html representation of the model will be displayed.
import pandas as pd
import daspi as dsp
df = dsp.load_dataset('painkillers-dissolution')
lm = dsp.LinearModel(
source=df,
target='dissolution',
factors=['employee', 'stirrer', 'brand', 'catalyst', 'water'],
covariates=['temperature', 'preparation'],
alpha=0.05,
order=3,
encode_categoricals=False
)
dsp.ParameterRelevanceCharts(lm).plot().stripes().label()
lm
To add again a factor to the model, use the include method.
For an automatic elimination of insignificant terms, use the
'recursive_elimination' method.
Notes
Always be careful when removing the intercept. If the intercept is missing, Patsy will automatically add all one-hot encoded levels for a categorical variable to compensate for this missing term. This will result in extremely high VIFs.
| RAISES | DESCRIPTION |
|---|---|
AssertionError:
|
If the given term is not in the model. |
include(parameter)
¶
Adds the given factor to the model by removing it from the
excluded set. Call fit to refit the model.
| PARAMETER | DESCRIPTION |
|---|---|
parameter
|
The factor name, the covariates name or the interaction of multiple factors to be added to the model.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
Self
|
The current instance of the model for more method chaining. |
Examples:
See eliminate method. After removing a term from the model,
you can add it again by using the include method.
recursive_elimination(rsquared_max=0.99, ensure_hierarchy=True, **kwds)
¶
Perform a recursive parameter elimination on the fitted model.
This function starts with the complete model and recursively eliminates parameters based on their p-values, until only significant parameters remain in the model. The function yields the goodness-of-fit metrics at each step of the elimination process.
| PARAMETER | DESCRIPTION |
|---|---|
rsquared_max
|
If given, the model must have a lower R^2 value than the given threshold, by default 0.99
TYPE:
|
ensure_hierarchy
|
Adds factors at the end to ensure model is hierarchical, by default True
TYPE:
|
**kwds
|
Additional keyword arguments for
DEFAULT:
|
| YIELDS | DESCRIPTION |
|---|---|
DataFrame
|
The goodness-of-fit metrics at each step of the recursive feature elimination. |
Examples:
Prepare a LinearModel instance, fit the model, automatically eliminate insignificant terms and plot the relevance of the parameters and the residuals. In the DataFrame df_gof you can see when which factor was removed and the current goodness-of-fit values can also be viewed. If you run the following code in a Jupyter Notebook, the plots and the html representation of the model will be displayed.
import pandas as pd
import daspi as dsp
df = dsp.load_dataset('painkillers-dissolution')
lm = dsp.LinearModel(
source=df,
target='dissolution',
factors=['employee', 'stirrer', 'brand', 'catalyst', 'water'],
covariates=['temperature', 'preparation'],
alpha=0.05,
order=3,
encode_categoricals=False
)
df_gof = pd.concat(list(lm.recursive_elimination()))
dsp.ParameterRelevanceCharts(lm).plot().stripes().label(info=True)
dsp.ResidualsCharts(lm).plot().stripes().label(info=True)
lm
anova(typ='', vif=False)
¶
Perform an analysis of variance (ANOVA) on the fitted model.
| PARAMETER | DESCRIPTION |
|---|---|
typ
|
The type of ANOVA to perform. Default is 'III', see notes for more informations about the types. - '' : If no or an invalid type is specified, Type-II is used if the model has no significant interactions. Otherwise, Type-III is used for hierarchical models and Type-I is used for non-hierarchical models. - 'I' : Type I sum of squares ANOVA. - 'II' : Type II sum of squares ANOVA. - 'III' : Type III sum of squares ANOVA.
TYPE:
|
vif
|
If True, variance inflation factors (VIF) are added to the anova table, by default False
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
The ANOVA table as DataFrame containing the following columns: - DF : Degrees of freedom for model terms. - SS : Sum of squares for model terms. - F : F statistic value for significance of adding model terms. - p : P-value for significance of adding model terms. - n2 : Eta-square as effect size (proportion of explained variance). - np2 : Partial eta-square as partial effect size. |
Notes
The ANOVA table provides information about the significance of each factor and interaction in the model. The type of ANOVA determines how the sum of squares is partitioned among the factors.
The SAS and also Minitab software uses Type III by default. This type is also the only one who gives us a SS and p-value for the Intercept. So Type-III is also used internaly for evaluating the least significant term. A discussion on which one to use can be found here: https://stats.stackexchange.com/a/93031
A nice conclusion about the differences between the types: - Typ-I: We choose the most "important" independent variable and it will receive the maximum amount of variation possible. - Typ-II: We ignore the shared variation: no interaction is assumed. If this is true, the Type II Sums of Squares are statistically more powerful. However if in reality there is an interaction effect, the model will be wrong and there will be a problem in the conclusions of the analysis. - Typ-III: If there is an interaction effect and we are looking for an “equal” split between the independent variables, Type-III should be used.
source: https://towardsdatascience.com/anovas-three-types-of-estimating-sums-of-squares-don-t-make-the-wrong-choice-91107c77a27a
Examples:
parameter_statistics(alpha=0.05, use_t=True)
¶
Calculate the parameter statistics for the fitted model.
| PARAMETER | DESCRIPTION |
|---|---|
alpha
|
The significance level for the confidence intervals, by default 0.05.
TYPE:
|
use_t
|
If True, use t-distribution for hypothesis testing and confidence intervals. If False, use normal distribution, by default True.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
A pandas DataFrame containing the parameter statistics for the fitted model. The DataFrame includes columns for the parameter estimates, standard errors, t-values (or z-values), and p-values. |
variance_inflation_factor(threshold=5)
¶
Calculate the variance inflation factor (VIF) and the generalized variance inflation factor (GVIF) for each predictor variable in the fitted model.
This function takes a regression model as input and returns a DataFrame containing the VIF, GVIF (= VIF^(1/2*dof)), threshold for GVIF, collinearity status and calculation kind for each predictor variable in the model. The VIF and GVIF are measures of multicollinearity, which can help identify variables that are highly correlated with each other.
| PARAMETER | DESCRIPTION |
|---|---|
threshold
|
The threshold for deciding whether a predictor is collinear. Common values are 5 and 10. By default 5.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
A DataFrame containing the VIF, GVIF, threshold, collinearity status and performed method for each predictor variable in the model. |
highest_parameters(factors_only=False)
¶
Determines all main and interaction parameters that do not occur in a higher interaction in this constellation.
| PARAMETER | DESCRIPTION |
|---|---|
factors_only
|
If true, intercept and interaction parameters are not returned. By default False.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
list[str]
|
list of highest parameters. |
has_insignificant_term(rsquared_max=0.99)
¶
Check if the fitted model has any insignificant terms.
| PARAMETER | DESCRIPTION |
|---|---|
rsquared_max
|
The maximum R^2 value that the model can have to be considered significant. If not provided, by default 0.99.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
bool
|
Returns True if the model has any insignificant terms, and False otherwise. |
has_significant_interactions()
¶
True if fitted model has significant interactions.
predict(xs)
¶
Predict y with given xs. Ensure that all non interactions are given in xs
| PARAMETER | DESCRIPTION |
|---|---|
xs
|
The values for which you want to predict. Make sure that all non interaction parameters are given in xs. If multiple values are to be predicted, provide a list of values for each factor level.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
A DataFrame containing the predicted values for the given values of the predictor variables. |
optimize(maximize=True, bounds=None)
¶
Optimize the prediction by optimizing the parameters.
| PARAMETER | DESCRIPTION |
|---|---|
maximize
|
Whether to maximize the prediction or minimize it.
TYPE:
|
bounds
|
Bounds for the parameters to optimize, default is None. - You can freeze a paramater by setting it to the desired value. For example, to fix the value of a parameter to 1, you can set bounds = {'param_name': 1}. - To keep an ordinal or metric parameter within a specific range, you can set bounds = {'param_name': (lower, upper)}. For example, to constrain a parameter to be between 0 and 1, you can set bounds = {'param_name': (0, 1)}. - In order to limit the selection of a nominal parameter only to a certain subset of the originally conained values, you can set bounds = {'param_name': (val1, val2, ...)}. For example, to limit the selection of a nominal parameter to the values 'A' and 'B', you can set bounds = {'param_name': ('A', 'B')}.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
dict[str, Any]
|
The optimized parameters. |
residual_data()
¶
Get the residual data from the fitted model.
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
The residual data containing the residuals, observation index, and predicted values. |
Examples:
import daspi as dsp
df = dsp.load_dataset('partial_factorial')
target = 'Yield'
factors = [c for c in df.columns if c != target]
lm = LinearModel(df, target, factors)
print(lm.residual_data())
Observation Residuals Prediction
0 1 9.250000e+00 46.75
1 2 2.000000e+00 51.00
2 3 -1.050000e+01 73.50
3 4 -2.500000e-01 65.25
4 5 -1.421085e-14 53.00
5 6 1.025000e+01 44.75
6 7 -2.500000e-01 67.25
7 8 -1.050000e+01 71.50
8 9 3.750000e+00 65.25
9 10 -1.200000e+01 57.00
10 11 -1.500000e+00 79.50
11 12 9.250000e+00 83.75
12 13 -1.000000e+01 59.00
13 14 -3.250000e+00 63.25
14 15 9.250000e+00 85.75
15 16 4.500000e+00 77.50
daspi.anova.model.GeneralizedLinearModel(source, target, factors, family='poisson', offset=None, alpha=0.05, order=1, fit_at_init=True)
¶
Bases: BaseHTMLReprModel
Generalized Linear Model for count data and proportions.
This class extends the functionality of LinearModel to handle non-normal response distributions, specifically: - Poisson distribution for count data (e.g., defect counts) - Negative Binomial for overdispersed count data - Binomial for proportions/binary data
Unlike standard ANOVA which assumes normally distributed errors, GLM uses maximum likelihood estimation and is appropriate for: - Count data (number of defects, failures, events) - Rate data (defects per unit) - Proportion data (pass/fail ratios) - Unbalanced designs with different sample sizes
| PARAMETER | DESCRIPTION |
|---|---|
source
|
Pandas DataFrame in long format containing the data.
TYPE:
|
target
|
Column name of the response variable (counts or proportions).
TYPE:
|
factors
|
Column names of predictor variables (categorical or continuous).
TYPE:
|
family
|
Distribution family: - 'poisson': For count data with mean ≈ variance - 'negbin': For overdispersed count data (variance > mean) - 'binomial': For proportion/binary data Default is 'poisson'.
TYPE:
|
offset
|
Column name for offset variable (e.g., exposure time, sample size). Used to model rates instead of counts. Default is None.
TYPE:
|
alpha
|
Significance level for hypothesis tests. Default is 0.05.
TYPE:
|
order
|
Maximum interaction order to include. Default is 1 (main effects only).
TYPE:
|
fit_at_init
|
Whether to fit the model immediately. Default is True.
TYPE:
|
Examples:
Analyze defect counts from an unbalanced factorial experiment:
import pandas as pd
import daspi as dsp
# Example data: defect counts for different factor combinations
data = {
'Factor_A': ['Low', 'Low', 'High', 'High', 'High'],
'Factor_B': ['X', 'Y', 'X', 'Y', 'X'],
'Defects': [12, 8, 5, 3, 6],
'Sample_Size': [100, 100, 100, 100, 100]
}
df = pd.DataFrame(data)
# Fit Poisson GLM
model = dsp.GLM(
source=df,
target='Defects',
factors=['Factor_A', 'Factor_B'],
family='poisson',
order=2 # Include interaction
)
# View results
print(model.summary())
print(model.deviance_check())
# Get factor effects
print(model.effects())
Check for overdispersion and switch to Negative Binomial if needed:
# Check dispersion
if model.dispersion > 1.5:
print("Overdispersion detected! Using Negative Binomial instead.")
model = dsp.GLM(
source=df,
target='Defects',
factors=['Factor_A', 'Factor_B'],
family='negbin',
order=2
)
Notes
When to use GLM instead of standard ANOVA: - Response variable is count data (discrete, non-negative integers) - Variance increases with the mean - Data is unbalanced (unequal sample sizes per group) - You want to model rates (e.g., defects per hour)
Interpreting coefficients: For Poisson/Negative Binomial models with log link: - exp(coef) gives the multiplicative effect on the mean - exp(coef) - 1 gives the % change in the mean
Checking model fit:
- Use deviance_check() to assess goodness of fit
- Dispersion ≈ 1 indicates good fit for Poisson
- Dispersion > 1.5 suggests overdispersion → use Negative Binomial
target = target
instance-attribute
¶
Column name of the response variable.
factors = factors
instance-attribute
¶
List of factor column names used in the model.
family = family
instance-attribute
¶
The distribution family for the GLM.
offset = offset
instance-attribute
¶
Column name of the offset variable, if any.
target_map = {target: 'y'}
instance-attribute
¶
Mapping of original target name to internal name ('y').
feature_map = {f: _f for f, _f in zip(factors, f_main_terms)}
instance-attribute
¶
Mapping of original factor names to internal names (x0, x1, ...).
main_term_map = {v: k for k, v in self.feature_map.items()}
instance-attribute
¶
Mapping of internal names back to original factor names.
data = source.rename(columns=self.feature_map | self.target_map | offset_map)[columns_to_keep].copy()
instance-attribute
¶
The data used for fitting the model, with renamed columns for internal processing.
print_formula = True
instance-attribute
¶
Whether to print the model formula.
fitted
property
¶
Whether the model is fitted (read-only).
model
property
¶
Get fitted GLM results. Calls fit method if not yet fitted.
alpha
property
writable
¶
Significance threshold for hypothesis tests.
formula
property
¶
Get the model formula (read-only).
deviance
property
¶
Deviance of the fitted model (goodness of fit measure).
dispersion
property
¶
Dispersion parameter: deviance / df_resid.
Values > 1 indicate overdispersion (variance > mean). For Poisson models, dispersion should be ≈ 1. If dispersion > 1.5, consider using Negative Binomial family.
fit(**kwds)
¶
Fit the Generalized Linear Model.
| PARAMETER | DESCRIPTION |
|---|---|
**kwds
|
Additional keyword arguments passed to GLM.fit()
DEFAULT:
|
| RETURNS | DESCRIPTION |
|---|---|
Self
|
The fitted model instance for method chaining. |
deviance_check()
¶
Check model fit using deviance statistics.
| RETURNS | DESCRIPTION |
|---|---|
dict[str, Any]
|
dictionary containing:
- 'deviance': Model deviance
- 'df_resid': Residual degrees of freedom |
summary()
¶
Generate a comprehensive summary of the fitted model.
parameter_statistics(alpha=None)
¶
Get parameter statistics table.
| PARAMETER | DESCRIPTION |
|---|---|
alpha
|
Significance level for confidence intervals. If None, uses self.alpha.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
Parameter statistics including coefficients, standard errors, z-values, p-values, and confidence intervals. |
effects()
¶
Calculate standardized effects (coef / std_err).
| RETURNS | DESCRIPTION |
|---|---|
Series
|
Standardized effects for each parameter. |
p_values()
¶
Get p-values for all parameters.
| RETURNS | DESCRIPTION |
|---|---|
Series
|
P-values from Wald tests for each parameter. |
predict(xs)
¶
Predict response for given factor values.
| PARAMETER | DESCRIPTION |
|---|---|
xs
|
Dictionary mapping factor names to values. Can be single values or lists for multiple predictions.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
Predictions with input values and predicted response. |
compare_groups()
¶
Compare predicted means across factor combinations.
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
Table showing mean predictions for each factor combination. |