Skip to content

cca_zoo.gp

Gaussian-process (GP) nonlinear CCA methods, fit by L-BFGS-B. Built entirely on scikit-learn's own GaussianProcessRegressor/RBF/ConstantKernel/KernelCenterer — no optional dependency required.


GaussianProcessCCA

GaussianProcessCCA(
    latent_dimensions: int = 1,
    center: bool = True,
    kernel: Kernel | list[Kernel | None] | None = None,
    alpha: float | list[float] = 0.01,
    n_inducing: int | list[int | None] | None = None,
    max_iter: int = 1000,
    tol: float = 1e-06,
    random_state: int = 0,
)

Bases: BaseModel

GaussianProcessCCA — nonlinear multiview CCA with Gaussian-process encoders.

Learns one nonlinear encoder \(f_i\) per view — a Gaussian process with a joint (non-additive) kernel over that view's raw feature vector — that jointly minimise the ridge-penalised Eckart-Young (EY) objective:

\[ \mathcal{L}_{EY} = -2 \operatorname{tr}(C) + \operatorname{tr}(V V) \]

where, for embeddings \(Z_i = f_i(X_i)\), \(C\) is the mean pairwise cross-covariance (including \(i = j\) terms) and \(V\) the mean auto-covariance across all views (see :mod:cca_zoo._utils._ey, the same shared EY-loss machinery used by :class:~cca_zoo.linear.gradient.CCAEY, :class:~cca_zoo.deep.DCCAEY, :class:~cca_zoo.tree.TreeCCA, :class:~cca_zoo.gam.GAMCCA, and :class:~cca_zoo.sparse.ElasticNetCCA). Writing \(f_i(x) = k(x, Z_i)^\top B_i\) for a fixed kernel \(k\) and fixed basis points \(Z_i\) (:class:_GpEncoder; every training row, or n_inducing of them selected via :func:sklearn.cluster.kmeans_plusplus) turns fitting \(B_i\) into an ordinary smooth optimisation problem with an exact, cheap analytic gradient — fit directly by L-BFGS-B (:func:scipy.optimize.minimize), the same algorithm :class:~sklearn.gaussian_process.GaussianProcessRegressor itself uses internally, just pointed at \(\mathcal{L}_{EY}\) (plus the RKHS-norm penalty \(B_i^\top K_{mm} B_i\)) instead of the negative log-marginal- likelihood it optimises kernel hyperparameters against. Every view's coefficients \(B_1, \dots, B_M\) are optimised jointly, in a single L-BFGS-B run over all of them concatenated, via :func:_lbfgsb_joint_objective — not one view at a time with the others held fixed: \(\mathcal{L}_{EY}\) already couples every view together, so a block Gauss-Seidel scheme (each view solved to convergence before moving to the next) needlessly repeats work and can settle into a worse joint optimum than optimising every view's coefficients simultaneously against the exact joint gradient. L-BFGS-B needs no explicit Hessian — just gradients.

A joint kernel is not restricted to a sum of univariate terms the way :class:~cca_zoo.gam.GAMCCA's additive splines are: it can represent a genuine interaction between two features of the same view directly.

Kernel hyperparameters (lengthscales, signal variance) are fixed — pass kernel explicitly, or tune it externally (e.g. with :class:~sklearn.model_selection.GridSearchCV, since this is an ordinary BaseEstimator) — there is no automatic marginal-likelihood search.

As a Bayesian model it still comes with calibrated predictive uncertainty for free: :meth:transform can return each latent component's posterior standard deviation alongside its mean. This does not depend on the fitted coefficients at all — a standard GP fact, since posterior variance only involves the kernel, the noise level, and the design points — so it is computed directly by an actual :class:~sklearn.gaussian_process.GaussianProcessRegressor fit with a placeholder target purely to reuse its variance formula (see :class:_GpEncoder).

Note

Exact and sparse inference are the same reduced-rank ("subset of regressors") construction, differing only in how many basis points \(Z_i\) are used (\(n\), i.e. every training row, for exact; n_inducing < \(n\), chosen by kmeans_plusplus, for sparse) — fitting costs \(O(n m^2 + m^3)\) for \(m\) basis points, linear in \(n\) once \(m \ll n\).

References

Rasmussen, C. E., & Williams, C. K. I. (2006). Gaussian Processes for Machine Learning. MIT Press.

Quiñonero-Candela, J., & Rasmussen, C. E. (2005). A Unifying View of Sparse Approximate Gaussian Process Regression. Journal of Machine Learning Research, 6, 1939-1959.

Chapman, J., Wells, L., & Lawry Aguila, A. (2024). Unconstrained Stochastic CCA: Unifying Multiview and Self-Supervised Learning. arXiv:2310.01012.

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent components. Must not exceed the number of features in any view. Default is 1.

1
center bool

Whether to subtract per-view column means before fitting. Default is True.

True
kernel Kernel | list[Kernel | None] | None

Kernel(s) used for each view. Either a single kernel (cloned per view -- its hyperparameters, e.g. an ARD length_scale array, must then be compatible with every view's feature count) or a list of one kernel per view. None entries (including a bare None, the default) fall back to ConstantKernel(1.0) * RBF(length_scale=np.ones(p)) for that view's own feature count p.

None
alpha float | list[float]

Ridge (RKHS-norm) penalty strength(s), also used as the noise level for the posterior-variance calculation. Either a single float applied to every view or a list of per-view floats. Default is 0.01.

0.01
n_inducing int | list[int | None] | None

Number of basis ("inducing") points. Either a single value or a list of per-view values. None (the default, either as a scalar or a per-view list entry) uses every training row (exact inference, appropriate up to a few thousand samples) for that view. For larger datasets, set this to a few hundred to make fitting scale as \(O(n \, \text{n\_inducing}^2)\) instead of \(O(n^3)\). Values at or above the number of training samples fall back to exact inference automatically.

None
max_iter int

Maximum number of L-BFGS-B iterations for the single, joint solve over every view's coefficients. Default is 1000.

1000
tol float

Convergence tolerance, passed to L-BFGS-B as ftol. Default is 1e-6.

1e-06
random_state int

Seed for the initial coefficients and (if n_inducing is set) for selecting inducing points.

0

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((100, 3))
>>> X2 = rng.standard_normal((100, 3))
>>> model = GaussianProcessCCA(latent_dimensions=1).fit([X1, X2])
>>> scores = model.transform([X1, X2])
>>> means, stds = model.transform([X1, X2], return_std=True)
>>> # For larger datasets, cap inference cost with inducing points:
>>> big_model = GaussianProcessCCA(latent_dimensions=1, n_inducing=200)

A different ridge penalty per view:

>>> model = GaussianProcessCCA(latent_dimensions=1, alpha=[0.01, 0.1]).fit(
...     [X1, X2]
... )
Source code in cca_zoo/gp/_gpcca.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    kernel: Kernel | list[Kernel | None] | None = None,
    alpha: float | list[float] = 0.01,
    n_inducing: int | list[int | None] | None = None,
    max_iter: int = 1000,
    tol: float = 1e-6,
    random_state: int = 0,
) -> None:
    super().__init__(latent_dimensions=latent_dimensions, center=center)
    self.kernel = kernel
    self.alpha = alpha
    self.n_inducing = n_inducing
    self.max_iter = max_iter
    self.tol = tol
    self.random_state = random_state

weights property

weights: list[ndarray]

Not implemented for GaussianProcessCCA.

Raises:

Type Description
NotFittedError

If fit has not been called.

NotImplementedError

GaussianProcessCCA encoders are Gaussian processes over the joint feature vector, not linear weight matrices, and have no per-feature decomposition analogous to :meth:~cca_zoo.gam.GAMCCA.shape_function (the kernel is not additive across features).

fit

fit(
    views: list[ArrayLike], y: None = None
) -> GaussianProcessCCA

Fit the GaussianProcessCCA model by L-BFGS-B on the EY loss.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of 2 or more arrays, each (n_samples, n_features_i).

required
y None

Ignored.

None

Returns:

Name Type Description
self GaussianProcessCCA

Fitted estimator.

Raises:

Type Description
ValueError

If fewer than 2 views are provided.

ValueError

If views have inconsistent numbers of samples.

Source code in cca_zoo/gp/_gpcca.py
def fit(self, views: list[ArrayLike], y: None = None) -> GaussianProcessCCA:
    """Fit the GaussianProcessCCA model by L-BFGS-B on the EY loss.

    Args:
        views: List of 2 or more arrays, each (n_samples, n_features_i).
        y: Ignored.

    Returns:
        self: Fitted estimator.

    Raises:
        ValueError: If fewer than 2 views are provided.
        ValueError: If views have inconsistent numbers of samples.
    """
    views_ = self._setup_fit(views)
    k = self.latent_dimensions

    kernel_ = perview_parameter("kernel", self.kernel, None, self.n_views_)
    alpha_ = perview_parameter("alpha", self.alpha, 0.01, self.n_views_)
    n_inducing_ = perview_parameter(
        "n_inducing", self.n_inducing, None, self.n_views_
    )

    encoders = [
        _GpEncoder(
            X,
            k,
            (
                clone(kern)
                if kern is not None
                else ConstantKernel(1.0) * RBF(length_scale=np.ones(X.shape[1]))
            ),
            a,
            n_ind,
            self.random_state,
        )
        for X, kern, a, n_ind in zip(views_, kernel_, alpha_, n_inducing_)
    ]

    bases = [enc.basis_ for enc in encoders]
    ridge_matrices = [enc.ridge_matrix_ for enc in encoders]

    rng = np.random.default_rng(self.random_state)
    coefficients0 = cheap_orthonormal_projection_weights(bases, k, None, rng)
    shapes = [c.shape for c in coefficients0]
    sizes = [c.size for c in coefficients0]
    x0 = np.concatenate([c.ravel() for c in coefficients0])

    result = minimize(
        _lbfgsb_joint_objective,
        x0,
        args=(bases, ridge_matrices, shapes, sizes, alpha_),
        jac=True,
        method="L-BFGS-B",
        options={"maxiter": self.max_iter, "ftol": self.tol},
    )

    coefficients = []
    offset = 0
    for shape, size in zip(shapes, sizes):
        coefficients.append(result.x[offset : offset + size].reshape(shape))
        offset += size

    representations = [basis @ coef for basis, coef in zip(bases, coefficients)]

    for enc, coef, rep in zip(encoders, coefficients, representations):
        enc.coef_ = coef
        enc._train_pred = rep

    self.encoders_: list[_GpEncoder] = encoders
    return self

transform

transform(
    views: list[ArrayLike], return_std: bool = False
) -> (
    list[np.ndarray]
    | tuple[list[np.ndarray], list[np.ndarray]]
)

Project views into the latent space using the fitted encoders.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of arrays, each (n_samples, n_features_i), matching the number of views passed to fit.

required
return_std bool

If True, also return each view's posterior standard deviation alongside its mean.

False

Returns:

Type Description
list[ndarray] | tuple[list[ndarray], list[ndarray]]

List of arrays, each (n_samples, latent_dimensions); or, if

list[ndarray] | tuple[list[ndarray], list[ndarray]]

return_std is True, a tuple (means, stds) of two such

list[ndarray] | tuple[list[ndarray], list[ndarray]]

lists.

Raises:

Type Description
NotFittedError

If fit has not been called.

ValueError

If fewer than 2 views are provided.

Source code in cca_zoo/gp/_gpcca.py
def transform(  # type: ignore[override]
    self, views: list[ArrayLike], return_std: bool = False
) -> list[np.ndarray] | tuple[list[np.ndarray], list[np.ndarray]]:
    """Project views into the latent space using the fitted encoders.

    Args:
        views: List of arrays, each (n_samples, n_features_i), matching
            the number of views passed to ``fit``.
        return_std: If True, also return each view's posterior standard
            deviation alongside its mean.

    Returns:
        List of arrays, each (n_samples, latent_dimensions); or, if
        ``return_std`` is True, a tuple ``(means, stds)`` of two such
        lists.

    Raises:
        sklearn.exceptions.NotFittedError: If ``fit`` has not been called.
        ValueError: If fewer than 2 views are provided.
    """
    check_is_fitted(self)
    validated = validate_views(views)
    centred = [v - m for v, m in zip(validated, self.means_)]
    if return_std:
        means = []
        stds = []
        for v, enc in zip(centred, self.encoders_):
            mean, std = enc.predict_new(v, return_std=True)
            means.append(mean)
            stds.append(std)
        return means, stds
    return [enc.predict_new(v) for v, enc in zip(centred, self.encoders_)]