Skip to content

cca_zoo.gam

Generalized-additive-model (GAM) nonlinear CCA methods, fit by a single joint trust-region Newton-CG solve (P-IRLS-style). Built entirely on scikit-learn's own SplineTransformer and scipy.optimize — no optional dependency required.


GAMCCA

GAMCCA(
    latent_dimensions: int = 1,
    center: bool = True,
    n_knots: int | list[int] = 5,
    alpha: float | list[float] = 0.1,
    max_iter: int = 100,
    tol: float = 1e-06,
    random_state: int = 0,
)

Bases: BaseModel

GAMCCA — nonlinear multiview CCA with generalized-additive-model encoders.

Learns one nonlinear encoder \(f_i\) per view — a generalized additive model (GAM), \(f_i(x) = \sum_j s_{ij}(x_{ij})\), summing one univariate B-spline term per input feature — that jointly minimise the elastic-net-penalised Eckart-Young (EY) objective:

\[ \mathcal{L}_{EY} = -2 \operatorname{tr}(C) + \operatorname{tr}(V V) \]

where, for embeddings \(Z_i = f_i(X_i)\), \(C\) is the mean pairwise cross-covariance (including \(i = j\) terms) and \(V\) the mean auto-covariance across all views (see :mod:cca_zoo._utils._ey, the same shared EY-loss machinery used by :class:~cca_zoo.linear.gradient.CCAEY, :class:~cca_zoo.deep.DCCAEY, :class:~cca_zoo.tree.TreeCCA, and :class:~cca_zoo.sparse.ElasticNetCCA). Writing \(f_i(x) = \sum_j s_{ij}(x_{ij})\) as \(\text{basis}_i(x) B_i\) for a fixed per-feature B-spline basis (:class:_GamEncoder, built by :class:~sklearn.preprocessing.SplineTransformer), fitting \(B_1, \dots, B_M\) is conceptually a P-IRLS problem — the same repeated-penalised- quadratic-solve structure mgcv itself uses to fit a GAM, applied directly to the EY loss rather than a per-observation likelihood — but rather than hand-rolling that solve (or even cycling over views Gauss-Seidel-style), every view's coefficients are updated jointly, in a single call: \(\mathcal{L}_{EY}\)'s own exact gradient (:func:~cca_zoo._utils._ey.ey_grad_z) and exact Hessian-vector product across the entire stacked parameter vector \((B_1, \dots, B_M)\) (:func:_gamcca_joint_hessp) are handed straight to :func:scipy.optimize.minimize's "trust-krylov" solver — a standard, off-the-shelf trust-region Newton-CG method — which performs all of its own outer Newton and inner Krylov iterations internally. \(\mathcal{L}_{EY}\) is not convex, so this Hessian is only guaranteed positive semi-definite near the loss's own fixed point; "trust-krylov" handles the indefinite case directly, and works from the Hessian-vector product alone, never forming or inverting the \(\left(\sum_i d_i k\right) \times \left(\sum_i d_i k\right)\) Hessian matrix explicitly. Each view's smoothing strength (alpha) is a fixed hyperparameter — there is no automatic smoothing-parameter search.

Because each latent component still decomposes exactly into one additive term per input feature, the fitted shape of any feature's contribution remains available directly via shape_function — the GAM analogue of :class:~cca_zoo.tree.TreeCCA's split-gain feature importance, but an exact curve rather than a single importance score.

Note

A GAM's additive structure assumes each feature contributes independently; it cannot represent a genuine interaction between two features of the same view (e.g. \(x_1 x_2\)) the way a multivariate tree split or a joint kernel can. If cross-view structure only shows up through such interactions, expect :class:~cca_zoo.tree.TreeCCA or :class:~cca_zoo.gp.GaussianProcessCCA to do better instead.

References

Wood, S. N. (2017). Generalized Additive Models: An Introduction with R (2nd ed.). Chapman and Hall/CRC.

Eilers, P. H., & Marx, B. D. (1996). Flexible smoothing with B-splines and penalties. Statistical Science, 11(2), 89-121.

Chapman, J., Wells, L., & Lawry Aguila, A. (2024). Unconstrained Stochastic CCA: Unifying Multiview and Self-Supervised Learning. arXiv:2310.01012.

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent components. Must not exceed the number of features in any view. Default is 1.

1
center bool

Whether to subtract per-view column means before fitting. Default is True.

True
n_knots int | list[int]

Number of knots per feature's B-spline term, passed straight through to sklearn.preprocessing.SplineTransformer( n_knots=...). Either a single value applied to every view or a list of per-view values. Default is 5.

5
alpha float | list[float]

Ridge (smoothing) penalty strength(s) applied to every spline coefficient. Either a single float applied to every view or a list of per-view floats. Default is 0.1.

0.1
max_iter int

Maximum number of outer Newton iterations in the single joint "trust-krylov" solve (scipy.optimize.minimize's own maxiter option). Default is 100.

100
tol float

Gradient-norm convergence tolerance for the joint solve (scipy.optimize.minimize's own gtol option). Default is 1e-6.

1e-06
random_state int

Seed for the initial coefficients.

0

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((200, 5))
>>> X2 = rng.standard_normal((200, 5))
>>> model = GAMCCA(latent_dimensions=2).fit([X1, X2])
>>> scores = model.transform([X1, X2])

A different number of knots and penalty per view:

>>> model = GAMCCA(latent_dimensions=2, n_knots=[5, 8], alpha=[0.1, 0.5]).fit(
...     [X1, X2]
... )
Source code in cca_zoo/gam/_gamcca.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    n_knots: int | list[int] = 5,
    alpha: float | list[float] = 0.1,
    max_iter: int = 100,
    tol: float = 1e-6,
    random_state: int = 0,
) -> None:
    super().__init__(latent_dimensions=latent_dimensions, center=center)
    self.n_knots = n_knots
    self.alpha = alpha
    self.max_iter = max_iter
    self.tol = tol
    self.random_state = random_state

weights property

weights: list[ndarray]

Not implemented for GAMCCA.

Raises:

Type Description
NotFittedError

If fit has not been called.

NotImplementedError

GAMCCA encoders are additive splines, not linear weight matrices. Use :meth:shape_function instead to inspect a feature's fitted contribution directly.

fit

fit(views: list[ArrayLike], y: None = None) -> GAMCCA

Fit the GAMCCA model by one joint trust-region Newton-CG solve.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of 2 or more arrays, each (n_samples, n_features_i).

required
y None

Ignored.

None

Returns:

Name Type Description
self GAMCCA

Fitted estimator.

Raises:

Type Description
ValueError

If fewer than 2 views are provided.

ValueError

If views have inconsistent numbers of samples.

Source code in cca_zoo/gam/_gamcca.py
def fit(self, views: list[ArrayLike], y: None = None) -> GAMCCA:
    """Fit the GAMCCA model by one joint trust-region Newton-CG solve.

    Args:
        views: List of 2 or more arrays, each (n_samples, n_features_i).
        y: Ignored.

    Returns:
        self: Fitted estimator.

    Raises:
        ValueError: If fewer than 2 views are provided.
        ValueError: If views have inconsistent numbers of samples.
    """
    views_ = self._setup_fit(views)
    k = self.latent_dimensions
    m = len(views_)
    n_knots_ = perview_parameter("n_knots", self.n_knots, 5, self.n_views_)
    alpha_ = perview_parameter("alpha", self.alpha, 0.1, self.n_views_)
    encoders = [_GamEncoder(X, k, nk) for X, nk in zip(views_, n_knots_)]
    bases = [enc.basis_ for enc in encoders]
    dims = [basis.shape[1] for basis in bases]
    grams = [basis.T @ basis for basis in bases]
    cross = [[bases[i].T @ bases[a] for a in range(m)] for i in range(m)]

    rng = np.random.default_rng(self.random_state)
    coefficients = cheap_orthonormal_projection_weights(bases, k, None, rng)
    x0 = _flatten(coefficients)

    result = minimize(
        _gamcca_joint_obj_grad,
        x0,
        args=(bases, grams, cross, alpha_, dims, k),
        jac=True,
        hessp=_gamcca_joint_hessp,
        method="trust-krylov",
        options={"maxiter": self.max_iter, "gtol": self.tol},
    )
    coefficients = _unflatten(result.x, dims, k)
    representations = [basis @ coef for basis, coef in zip(bases, coefficients)]

    for enc, coef, rep in zip(encoders, coefficients, representations):
        enc.coef_ = coef
        enc._train_pred = rep

    self.encoders_: list[_GamEncoder] = encoders
    return self

transform

transform(views: list[ArrayLike]) -> list[np.ndarray]

Project views into the latent space using the fitted encoders.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of arrays, each (n_samples, n_features_i), matching the number of views passed to fit.

required

Returns:

Type Description
list[ndarray]

List of arrays, each (n_samples, latent_dimensions).

Raises:

Type Description
NotFittedError

If fit has not been called.

ValueError

If fewer than 2 views are provided.

Source code in cca_zoo/gam/_gamcca.py
def transform(self, views: list[ArrayLike]) -> list[np.ndarray]:
    """Project views into the latent space using the fitted encoders.

    Args:
        views: List of arrays, each (n_samples, n_features_i), matching
            the number of views passed to ``fit``.

    Returns:
        List of arrays, each (n_samples, latent_dimensions).

    Raises:
        sklearn.exceptions.NotFittedError: If ``fit`` has not been called.
        ValueError: If fewer than 2 views are provided.
    """
    check_is_fitted(self)
    validated = validate_views(views)
    centred = [v - m for v, m in zip(validated, self.means_)]
    return [enc.predict_new(v) for v, enc in zip(centred, self.encoders_)]

shape_function

shape_function(
    view: int, feature: int, x: ArrayLike
) -> np.ndarray

Evaluate one feature's fitted additive term \(s_j(x_j)\).

Because GAMCCA's encoder is additive across features, each term can be inspected in isolation — the direct GAM analogue of :class:~cca_zoo.tree.TreeCCA's split-gain feature importance, but an exact, shape-preserving curve rather than a single importance score.

Parameters:

Name Type Description Default
view int

Index of the view.

required
feature int

Index of the feature within that view (raw, i.e. un-centred column order).

required
x ArrayLike

Raw (un-centred) values for that feature at which to evaluate the term, shape (n,).

required

Returns:

Type Description
ndarray

Array of shape (n, latent_dimensions): that feature's

ndarray

contribution alone, for every latent component.

Raises:

Type Description
NotFittedError

If fit has not been called.

Source code in cca_zoo/gam/_gamcca.py
def shape_function(self, view: int, feature: int, x: ArrayLike) -> np.ndarray:
    r"""Evaluate one feature's fitted additive term $s_j(x_j)$.

    Because GAMCCA's encoder is additive across features, each term can
    be inspected in isolation — the direct GAM analogue of
    :class:`~cca_zoo.tree.TreeCCA`'s split-gain feature importance, but
    an exact, shape-preserving curve rather than a single importance
    score.

    Args:
        view: Index of the view.
        feature: Index of the feature within that view (raw, i.e.
            un-centred column order).
        x: Raw (un-centred) values for that feature at which to evaluate
            the term, shape (n,).

    Returns:
        Array of shape (n, latent_dimensions): that feature's
        contribution alone, for every latent component.

    Raises:
        sklearn.exceptions.NotFittedError: If ``fit`` has not been called.
    """
    check_is_fitted(self)
    x_arr = np.asarray(x, dtype=float) - self.means_[view][feature]
    return self.encoders_[view].feature_term(feature, x_arr)