cca_zoo.gp¶
Gaussian-process (GP) nonlinear CCA methods, fit by L-BFGS-B. Built entirely on scikit-learn's own
GaussianProcessRegressor/RBF/ConstantKernel/KernelCenterer — no optional dependency
required.
GaussianProcessCCA ¶
GaussianProcessCCA(
latent_dimensions: int = 1,
center: bool = True,
kernel: Kernel | list[Kernel | None] | None = None,
alpha: float | list[float] = 0.01,
n_inducing: int | list[int | None] | None = None,
max_iter: int = 1000,
tol: float = 1e-06,
random_state: int = 0,
)
Bases: BaseModel
GaussianProcessCCA — nonlinear multiview CCA with Gaussian-process encoders.
Learns one nonlinear encoder \(f_i\) per view — a Gaussian process with a joint (non-additive) kernel over that view's raw feature vector — that jointly minimise the ridge-penalised Eckart-Young (EY) objective:
where, for embeddings \(Z_i = f_i(X_i)\), \(C\) is the mean pairwise
cross-covariance (including \(i = j\) terms) and \(V\) the mean
auto-covariance across all views (see :mod:cca_zoo._utils._ey, the
same shared EY-loss machinery used by
:class:~cca_zoo.linear.gradient.CCAEY, :class:~cca_zoo.deep.DCCAEY,
:class:~cca_zoo.tree.TreeCCA, :class:~cca_zoo.gam.GAMCCA, and
:class:~cca_zoo.sparse.ElasticNetCCA). Writing \(f_i(x) =
k(x, Z_i)^\top B_i\) for a fixed kernel \(k\) and fixed basis points \(Z_i\)
(:class:_GpEncoder; every training row, or n_inducing of them
selected via :func:sklearn.cluster.kmeans_plusplus) turns fitting
\(B_i\) into an ordinary smooth optimisation problem with an exact,
cheap analytic gradient — fit directly by L-BFGS-B
(:func:scipy.optimize.minimize), the same algorithm
:class:~sklearn.gaussian_process.GaussianProcessRegressor itself uses
internally, just pointed at \(\mathcal{L}_{EY}\) (plus the RKHS-norm
penalty \(B_i^\top K_{mm} B_i\)) instead of the negative log-marginal-
likelihood it optimises kernel hyperparameters against. Every view's
coefficients \(B_1, \dots, B_M\) are optimised jointly, in a single
L-BFGS-B run over all of them concatenated, via
:func:_lbfgsb_joint_objective — not one view at a time with the
others held fixed: \(\mathcal{L}_{EY}\) already couples every view
together, so a block Gauss-Seidel scheme (each view solved to
convergence before moving to the next) needlessly repeats work and can
settle into a worse joint optimum than optimising every view's
coefficients simultaneously against the exact joint gradient.
L-BFGS-B needs no explicit Hessian — just gradients.
A joint kernel is not restricted to a sum of univariate terms the way
:class:~cca_zoo.gam.GAMCCA's additive splines are: it can represent a
genuine interaction between two features of the same view directly.
Kernel hyperparameters (lengthscales, signal variance) are fixed —
pass kernel explicitly, or tune it externally (e.g. with
:class:~sklearn.model_selection.GridSearchCV, since this is an
ordinary BaseEstimator) — there is no automatic marginal-likelihood
search.
As a Bayesian model it still comes with calibrated predictive
uncertainty for free: :meth:transform can return each latent
component's posterior standard deviation alongside its mean. This does
not depend on the fitted coefficients at all — a standard GP fact, since
posterior variance only involves the kernel, the noise level, and the
design points — so it is computed directly by an actual
:class:~sklearn.gaussian_process.GaussianProcessRegressor fit with a
placeholder target purely to reuse its variance formula (see
:class:_GpEncoder).
Note
Exact and sparse inference are the same reduced-rank
("subset of regressors") construction, differing only in how many
basis points \(Z_i\) are used (\(n\), i.e. every training row, for
exact; n_inducing < \(n\), chosen by kmeans_plusplus, for
sparse) — fitting costs \(O(n m^2 + m^3)\) for \(m\) basis points,
linear in \(n\) once \(m \ll n\).
References
Rasmussen, C. E., & Williams, C. K. I. (2006). Gaussian Processes for Machine Learning. MIT Press.
Quiñonero-Candela, J., & Rasmussen, C. E. (2005). A Unifying View of Sparse Approximate Gaussian Process Regression. Journal of Machine Learning Research, 6, 1939-1959.
Chapman, J., Wells, L., & Lawry Aguila, A. (2024). Unconstrained Stochastic CCA: Unifying Multiview and Self-Supervised Learning. arXiv:2310.01012.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent components. Must not exceed the number of features in any view. Default is 1. |
1
|
center
|
bool
|
Whether to subtract per-view column means before fitting. Default is True. |
True
|
kernel
|
Kernel | list[Kernel | None] | None
|
Kernel(s) used for each view. Either a single kernel
(cloned per view -- its hyperparameters, e.g. an ARD
|
None
|
alpha
|
float | list[float]
|
Ridge (RKHS-norm) penalty strength(s), also used as the noise level for the posterior-variance calculation. Either a single float applied to every view or a list of per-view floats. Default is 0.01. |
0.01
|
n_inducing
|
int | list[int | None] | None
|
Number of basis ("inducing") points. Either a single
value or a list of per-view values. |
None
|
max_iter
|
int
|
Maximum number of L-BFGS-B iterations for the single, joint solve over every view's coefficients. Default is 1000. |
1000
|
tol
|
float
|
Convergence tolerance, passed to L-BFGS-B as |
1e-06
|
random_state
|
int
|
Seed for the initial coefficients and (if
|
0
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((100, 3))
>>> X2 = rng.standard_normal((100, 3))
>>> model = GaussianProcessCCA(latent_dimensions=1).fit([X1, X2])
>>> scores = model.transform([X1, X2])
>>> means, stds = model.transform([X1, X2], return_std=True)
>>> # For larger datasets, cap inference cost with inducing points:
>>> big_model = GaussianProcessCCA(latent_dimensions=1, n_inducing=200)
A different ridge penalty per view:
Source code in cca_zoo/gp/_gpcca.py
weights
property
¶
Not implemented for GaussianProcessCCA.
Raises:
| Type | Description |
|---|---|
NotFittedError
|
If |
NotImplementedError
|
GaussianProcessCCA encoders are Gaussian processes over
the joint feature vector, not linear weight matrices, and
have no per-feature decomposition analogous to
:meth: |
fit ¶
Fit the GaussianProcessCCA model by L-BFGS-B on the EY loss.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of 2 or more arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
GaussianProcessCCA
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If fewer than 2 views are provided. |
ValueError
|
If views have inconsistent numbers of samples. |
Source code in cca_zoo/gp/_gpcca.py
transform ¶
transform(
views: list[ArrayLike], return_std: bool = False
) -> (
list[np.ndarray]
| tuple[list[np.ndarray], list[np.ndarray]]
)
Project views into the latent space using the fitted encoders.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each (n_samples, n_features_i), matching
the number of views passed to |
required |
return_std
|
bool
|
If True, also return each view's posterior standard deviation alongside its mean. |
False
|
Returns:
| Type | Description |
|---|---|
list[ndarray] | tuple[list[ndarray], list[ndarray]]
|
List of arrays, each (n_samples, latent_dimensions); or, if |
list[ndarray] | tuple[list[ndarray], list[ndarray]]
|
|
list[ndarray] | tuple[list[ndarray], list[ndarray]]
|
lists. |
Raises:
| Type | Description |
|---|---|
NotFittedError
|
If |
ValueError
|
If fewer than 2 views are provided. |