cca_zoo.linear¶
Linear CCA methods. All classes are sklearn.base.BaseEstimator subclasses.
Sparse/regularised iterative methods live in cca_zoo.sparse;
mini-batch methods live in cca_zoo.stochastic.
Base class¶
BaseModel ¶
Bases: BaseEstimator, ABC
Abstract base class for all multiview CCA models.
Subclasses must implement :meth:fit. All other public methods
(transform, inverse_transform, fit_transform, score,
pairwise_correlations, average_pairwise_correlations,
get_factor_loadings, predict) are provided here using the
weights_ attribute set by fit.
This class inherits from :class:sklearn.base.BaseEstimator so that
get_params / set_params round-trip correctly and sklearn model
selection utilities work out of the box.
Constructor parameters are validated with sklearn's
_parameter_constraints mechanism (see :meth:_setup_fit).
Subclasses that add their own constructor parameters may extend
_parameter_constraints by merging in BaseModel._parameter_constraints;
parameters with no declared constraint are left unvalidated, so this is
always safe to skip.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions to fit. Default is 1. |
1
|
center
|
bool
|
Whether to subtract per-view column means before fitting.
The means are stored in |
True
|
weights
property
¶
Weight matrices post-fit, one per view.
Shape is (n_features_i, latent_dimensions) for each view.
Raises:
| Type | Description |
|---|---|
NotFittedError
|
If |
fit
abstractmethod
¶
Fit the model to multiview data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each of shape (n_samples, n_features_i). All arrays must have the same number of rows. |
required |
y
|
None
|
Ignored. Present for scikit-learn API compatibility. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
BaseModel
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If fewer than 2 views are provided. |
ValueError
|
If views have inconsistent numbers of samples. |
transform ¶
Project views into the latent space using the fitted weights.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each of shape (n_samples, n_features_i). |
required |
Returns:
| Type | Description |
|---|---|
list[ndarray]
|
List of arrays, each of shape (n_samples, latent_dimensions). |
Raises:
| Type | Description |
|---|---|
NotFittedError
|
If |
inverse_transform ¶
Approximately invert transform, mapping latent scores back to each view.
Each view is reconstructed from that same view's own latent
score only, via a per-view loading matrix fit by least squares at
fit time: the regression of that view's centred training data onto
that view's own training latent score. This makes
inverse_transform(transform(views)) an approximate round trip
of views (exact wherever latent_dimensions and each view's
own weights span it exactly), mirroring
:meth:sklearn.decomposition.PCA.inverse_transform.
This is a different operation from :meth:predict: predict
combines the observed views' scores into one shared consensus
estimate to reconstruct views you don't have, which is only
possible once at least one other view actually is observed.
inverse_transform never mixes information across views — it
needs a view's own score to reconstruct that same view, so it
cannot be used to impute a view you never transformed in the first
place; use :meth:predict for that.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
scores
|
list[ArrayLike]
|
List of length |
required |
Returns:
| Type | Description |
|---|---|
list[ndarray]
|
List of length |
(n_samples, n_features_i)
|
the reconstructed view. |
Raises:
| Type | Description |
|---|---|
NotFittedError
|
If |
ValueError
|
If |
Examples:
>>> import numpy as np
>>> from cca_zoo.linear import CCA
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = CCA(latent_dimensions=2).fit([X1, X2])
>>> scores = model.transform([X1, X2])
>>> X1_approx, X2_approx = model.inverse_transform(scores)
>>> X1_approx.shape
(50, 10)
predict ¶
Reconstruct every view from whichever views are observed.
transform maps data to the shared latent space; predict maps
back the other way, from the latent space to each view's original
feature space. Pass None for any view you want reconstructed —
typically one you don't have, but you can also ask for a view you
did supply, as a diagnostic (its own self-reconstruction).
The shared latent score is estimated as the mean, over the observed
views only, of that view's own projection ((view - mean) @
weights). Each requested view is then reconstructed as that score
against a per-view loading matrix fitted by least squares at fit
time: the regression of that view's centred training data onto the
training data's own shared latent score.
This regression-based reconstruction is deliberate rather than the
simpler score @ weights.T: for CCA (unlike PLS), that simpler
formula is only a correct inverse of transform when the data
happens to be pre-whitened, since the true forward map needs an
extra view-covariance factor that isn't recovered from weights_
alone (see the discussion on
https://github.com/jameschapman19/cca_zoo/issues/182). Fitting the
reconstruction directly against training data sidesteps that
entirely, at the cost of a fitted model retaining a reference to
its own (centred) training views.
See also :meth:inverse_transform, which reconstructs a view from
that same view's own score (no cross-view imputation) — the
appropriate choice when you already have every view's scores and
just want to invert transform.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike | None]
|
List of length |
required |
Returns:
| Type | Description |
|---|---|
list[ndarray]
|
List of length |
list[ndarray]
|
of shape (n_samples, n_features_i) with the same |
list[ndarray]
|
as the observed view(s) passed in. |
Raises:
| Type | Description |
|---|---|
NotFittedError
|
If |
ValueError
|
If |
Examples:
fit_transform ¶
Fit and then transform the training data.
Equivalent to self.fit(views).transform(views) but may be more
efficient for some subclasses.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each of shape (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Type | Description |
|---|---|
list[ndarray]
|
List of arrays, each of shape (n_samples, latent_dimensions). |
score ¶
Return average pairwise canonical correlations for each dimension.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each of shape (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Type | Description |
|---|---|
ndarray
|
Array of shape |
ndarray
|
pairwise correlation for each canonical dimension. |
pairwise_correlations ¶
Compute the full pairwise correlation matrix per latent dimension.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each of shape (n_samples, n_features_i). |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
Array of shape |
ndarray
|
entry |
ndarray
|
d-th canonical variate of view i and view j. |
average_pairwise_correlations ¶
Return the mean off-diagonal pairwise correlation per dimension.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each of shape (n_samples, n_features_i). |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
Array of shape |
ndarray
|
off-diagonal pairwise correlation for each canonical dimension. |
get_factor_loadings ¶
Compute canonical factor loadings for each view.
A loading is the Pearson correlation between an original feature and a canonical variate. Loadings indicate which original variables drive each canonical direction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each of shape (n_samples, n_features_i). |
required |
Returns:
| Type | Description |
|---|---|
list[ndarray]
|
List of arrays, each of shape (n_features_i, latent_dimensions), |
list[ndarray]
|
where entry |
list[ndarray]
|
view i and the d-th canonical variate of view i. |
Two-view exact methods¶
CCA ¶
Bases: rCCA
Canonical Correlation Analysis.
Finds the pair of linear projections that maximise the Pearson correlation between two views subject to unit within-view variance constraints:
This is a special case of :class:rCCA with c=0. The solution uses
PCA whitening followed by an SVD of the cross-covariance matrix, which is
numerically stable even for high-dimensional views.
References
Hotelling, H. (1936). Relations between two sets of variates. Biometrika, 28(3/4), 321–377.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means before fitting. Default True. |
True
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = CCA(latent_dimensions=2).fit([X1, X2])
>>> corrs = model.score([X1, X2])
Source code in cca_zoo/linear/_cca.py
fit ¶
Fit the CCA model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of exactly two arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
CCA
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the number of views is not exactly 2. |
ValueError
|
If views have inconsistent numbers of samples. |
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = CCA(latent_dimensions=2).fit([X1, X2])
Source code in cca_zoo/linear/_cca.py
rCCA ¶
Bases: BaseModel
Regularised Canonical Correlation Analysis (canonical ridge).
Finds the pair of linear projections of two views that maximise their correlation subject to regularised within-view variance constraints:
The solution is found by whitening each view with its regularised covariance matrix and computing the SVD of the resulting cross-covariance.
:class:CCA (c=0) and :class:PLS (c=1) are special cases.
References
Vinod, H. D. (1976). Canonical ridge and econometrics of joint production. Journal of Econometrics, 4(2), 147–166.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means before fitting. Default True. |
True
|
c
|
float | list[float]
|
Ridge regularisation parameter(s) in |
0.0
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = rCCA(latent_dimensions=2, c=0.1).fit([X1, X2])
>>> scores = model.transform([X1, X2])
Source code in cca_zoo/linear/_rcca.py
fit ¶
Fit the rCCA model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of exactly two arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
rCCA
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the number of views is not exactly 2. |
ValueError
|
If views have inconsistent numbers of samples. |
Source code in cca_zoo/linear/_rcca.py
PLS ¶
Bases: rCCA
Partial Least Squares (two-view).
Finds the pair of unit-norm weight vectors that maximise the covariance between the projected views:
This is equivalent to the truncated SVD of the sample cross-covariance
matrix \(X_1^\top X_2 / (n - 1)\), and corresponds to :class:rCCA
with c=1.
References
Wold, H. (1975). Soft modelling by latent variables: the nonlinear iterative partial least squares (NIPALS) approach. Perspectives in Probability and Statistics, 117–142.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means before fitting. Default True. |
True
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = PLS(latent_dimensions=2).fit([X1, X2])
>>> scores = model.transform([X1, X2])
Source code in cca_zoo/linear/_pls.py
fit ¶
Fit the PLS model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of exactly two arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
PLS
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the number of views is not exactly 2. |
ValueError
|
If views have inconsistent numbers of samples. |
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = PLS(latent_dimensions=2).fit([X1, X2])
Source code in cca_zoo/linear/_pls.py
Multiview methods¶
MCCA ¶
MCCA(
latent_dimensions: int = 1,
center: bool = True,
c: float | list[float] = 0.0,
pca: bool = True,
eps: float = 1e-06,
)
Bases: BaseModel
Multiset Canonical Correlation Analysis.
Finds linear projections of multiple (>=2) views that maximise the sum of
pairwise cross-view covariances subject to within-view variance constraints.
A ridge regularisation parameter c controls the trade-off between
correlation and variance explained.
The primal objective is:
This is solved as a generalised eigenvalue problem:
where \(A\) is the between-view block covariance matrix and \(B\) is the block-diagonal regularised within-view covariance matrix.
When pca=True (default), each view is first reduced to its principal
components, which makes the problem numerically stable for
high-dimensional data and allows an efficient closed-form \(B\).
References
Kettenring, J. R. (1971). Canonical analysis of several sets of variables. Biometrika, 58(3), 433–451.
Vinod, H. D. (1976). Canonical ridge and econometrics of joint production. Journal of Econometrics, 4(2), 147–166.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means before fitting. Default True. |
True
|
c
|
float | list[float]
|
Ridge regularisation parameter(s). Either a single float applied
to all views or a list of per-view floats in |
0.0
|
pca
|
bool
|
Whether to apply full PCA whitening as a pre-processing step before solving the eigenvalue problem. Highly recommended for high-dimensional data. Default is True. |
True
|
eps
|
float
|
Small constant added to the eigenvalues of B to ensure positive definiteness. Default is 1e-6. |
1e-06
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> X3 = rng.standard_normal((50, 6))
>>> model = MCCA(latent_dimensions=2).fit([X1, X2, X3])
>>> scores = model.transform([X1, X2, X3])
Source code in cca_zoo/linear/_mcca.py
fit ¶
Fit the MCCA model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
MCCA
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If fewer than 2 views are provided. |
ValueError
|
If views have inconsistent numbers of samples. |
Source code in cca_zoo/linear/_mcca.py
GCCA ¶
GCCA(
latent_dimensions: int = 1,
center: bool = True,
c: float | list[float] = 0.0,
view_weights: list[float] | None = None,
eps: float = 1e-06,
)
Bases: BaseModel
Generalised Canonical Correlation Analysis.
Finds linear projections of multiple (>=2) views that maximise their joint correlation with a shared auxiliary latent vector:
The solution is obtained by constructing the weighted projection matrix:
and computing its top-k eigenvectors \(V\), then recovering the per-view weights as \(\mathbf{w}_i = X_i^+ V\).
References
Tenenhaus, A., & Tenenhaus, M. (2011). Regularized generalized canonical correlation analysis. Psychometrika, 76(2), 257–284.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means before fitting. Default True. |
True
|
c
|
float | list[float]
|
Ridge regularisation parameter(s) in |
0.0
|
view_weights
|
list[float] | None
|
Per-view weights \(\mu_i\) in the GCCA objective. Default is equal weights (1 for all views). |
None
|
eps
|
float
|
Regularisation floor for within-view matrices. Default is 1e-6. |
1e-06
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> X3 = rng.standard_normal((50, 6))
>>> model = GCCA(latent_dimensions=2).fit([X1, X2, X3])
>>> scores = model.transform([X1, X2, X3])
Source code in cca_zoo/linear/_gcca.py
fit ¶
Fit the GCCA model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
GCCA
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If fewer than 2 views are provided. |
ValueError
|
If views have inconsistent numbers of samples. |
Source code in cca_zoo/linear/_gcca.py
TCCA ¶
TCCA(
latent_dimensions: int = 1,
center: bool = True,
c: float | list[float] = 0.0,
eps: float = 1e-06,
random_state: int | None = None,
)
Bases: BaseModel
Tensor Canonical Correlation Analysis.
Extends CCA to more than two views by exploiting higher-order cross-view correlations via a tensor product structure. The method constructs the order-M cross-moment tensor:
where \(\tilde{X}_j = X_j \Sigma_j^{-1/2}\) are the whitened views, and then decomposes \(\mathcal{M}\) using PARAFAC to recover the canonical directions.
References
Kim, T.-K., Wong, S.-F., & Cipolla, R. (2007). Tensor canonical correlation analysis for action classification. CVPR 2007. IEEE.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means before fitting. Default True. |
True
|
c
|
float | list[float]
|
Ridge regularisation in |
0.0
|
eps
|
float
|
Regularisation floor for within-view covariance matrices. |
1e-06
|
random_state
|
int | None
|
Seed for reproducibility (passed to PARAFAC). |
None
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 5))
>>> X2 = rng.standard_normal((50, 5))
>>> X3 = rng.standard_normal((50, 5))
>>> model = TCCA(latent_dimensions=2, random_state=0).fit([X1, X2, X3])
>>> scores = model.transform([X1, X2, X3])
Source code in cca_zoo/linear/_tcca.py
fit ¶
Fit the TCCA model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
TCCA
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If fewer than 2 views are provided. |
ValueError
|
If views have inconsistent numbers of samples. |
Source code in cca_zoo/linear/_tcca.py
Confound-adjusted / structured methods¶
PartialCCA ¶
PartialCCA(
latent_dimensions: int = 1,
center: bool = True,
c: float | list[float] = 0.0,
eps: float = 1e-06,
)
Bases: MCCA
Partial Canonical Correlation Analysis.
Extends CCA to account for confounding variables partials that may
drive the correlation between views. Each view is first deconfounded by
regressing out partials via least squares, and (ridge-regularised)
CCA is then applied to the residuals, subject to the additional
constraint that canonical weights are orthogonal to the confounds:
References
Rao, B. R. (1969). Partial canonical correlations. Trabajos de Estadistica y de Investigacion Operativa, 20(2-3), 211-219.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means before fitting. Default True. |
True
|
c
|
float | list[float]
|
Ridge regularisation parameter(s) applied to the deconfounded views. Either a scalar or a per-view list. Default is 0. |
0.0
|
eps
|
float
|
Small constant added to the eigenvalues of B to ensure positive definiteness. Default is 1e-6. |
1e-06
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> Z = rng.standard_normal((50, 3))
>>> model = PartialCCA(latent_dimensions=2).fit([X1, X2], partials=Z)
>>> scores = model.transform([X1, X2], partials=Z)
Source code in cca_zoo/linear/_partialcca.py
fit ¶
Fit the Partial CCA model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
partials
|
ArrayLike | None
|
Confound array of shape (n_samples, n_confounds) to regress out of each view before fitting CCA. Required. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
PartialCCA
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in cca_zoo/linear/_partialcca.py
transform ¶
Project views into the latent space, optionally removing confounds.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each (n_samples, n_features_i). |
required |
partials
|
ArrayLike | None
|
Confound array matching the one used at fit time. If
omitted, no deconfounding is applied (falls back to a plain
linear projection), which keeps |
None
|
Returns:
| Type | Description |
|---|---|
list[ndarray]
|
List of arrays, each (n_samples, latent_dimensions). |
Source code in cca_zoo/linear/_partialcca.py
fit_transform ¶
fit_transform(
views: list[ArrayLike],
y: None = None,
partials: ArrayLike | None = None,
) -> list[np.ndarray]
Fit and then transform the training data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
partials
|
ArrayLike | None
|
Confound array, passed through to both |
None
|
Returns:
| Type | Description |
|---|---|
list[ndarray]
|
List of arrays, each (n_samples, latent_dimensions). |
Source code in cca_zoo/linear/_partialcca.py
GRCCA ¶
GRCCA(
latent_dimensions: int = 1,
center: bool = True,
c: float | list[float] = 0.0,
mu: float | list[float] = 0.0,
eps: float = 1e-06,
)
Bases: MCCA
Group Regularised Canonical Correlation Analysis.
Extends :class:MCCA with structured ridge regularisation that shrinks
within-group feature weights toward a shared group-level effect. Each
view's features are partitioned into groups via feature_groups; the
per-view c parameter controls shrinkage of within-group deviations
and mu controls the weighting of the group-level effect.
Each view is internally augmented with group-mean features before
solving the generalised eigenvalue problem, then the resulting weights
are algebraically collapsed back to the original feature space, so
transform operates directly on the un-augmented views.
References
Tuzhilina, E., Tozzi, L., & Hastie, T. (2021). Canonical correlation analysis in high dimensions with structured regularization. Statistical Modelling.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means before fitting. Default True. |
True
|
c
|
float | list[float]
|
Ridge regularisation parameter(s) controlling within-group
shrinkage. Either a scalar applied to all views or a per-view
list, each in |
0.0
|
mu
|
float | list[float]
|
Regularisation parameter(s) controlling the group-level effect scale. Either a scalar or a per-view list. Default is 0. |
0.0
|
eps
|
float
|
Small constant added to the eigenvalues of B to ensure positive definiteness. Default is 1e-6. |
1e-06
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> groups1 = rng.integers(0, 3, size=10)
>>> groups2 = rng.integers(0, 3, size=8)
>>> model = GRCCA(latent_dimensions=2, c=0.5).fit(
... [X1, X2], feature_groups=[groups1, groups2]
... )
>>> scores = model.transform([X1, X2])
Source code in cca_zoo/linear/_grcca.py
fit ¶
fit(
views: list[ArrayLike],
y: None = None,
feature_groups: list[ndarray] | None = None,
) -> GRCCA
Fit the GRCCA model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
feature_groups
|
list[ndarray] | None
|
List of integer group-label arrays, one per
view, each of shape (n_features_i,). Required for
meaningful grouping whenever the corresponding per-view
|
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
GRCCA
|
Fitted estimator. |
Source code in cca_zoo/linear/_grcca.py
Reduced-rank regression methods¶
CCAR3 ¶
CCAR3(
latent_dimensions: int = 1,
center: bool = True,
lambda_: float = 0.0,
highdim: bool = True,
ledoit_wolf: bool = True,
max_iter: int = 10000,
tol: float = 0.0001,
eps: float = 1e-08,
)
Bases: BaseModel
Canonical Correlation Analysis via Reduced Rank Regression.
Recasts two-view CCA as a reduced-rank regression: Y is first
whitened by its (optionally Ledoit-Wolf shrunk) covariance,
and a coefficient matrix \(B\) relating \(X\) to \(\tilde{Y}\) is estimated.
In the low-dimensional regime (highdim=False) this has the closed
form \(B = \Sigma_X^{-1} X^\top \tilde{Y} / n\) (an ordinary reduced-rank
regression, distinct from the classical CCA eigenproblem — the two
agree only when \(\Sigma_X\) is close to isotropic). In the
high-dimensional regime (highdim=True, the default), \(B\) is
instead estimated by a row-wise group-lasso-penalised regression,
which is exactly the problem :class:~sklearn.linear_model.MultiTaskLasso
solves (at alpha = lambda_ / 2, to match sklearn's own \(\frac{1}{2n}\)
loss convention), so it's solved by delegating to that estimator's
coordinate-descent solver rather than a hand-rolled one. This drives
whole rows of \(B\) (whole \(X\) features) to zero, giving a
sparse-in-\(X\) solution well-suited to \(p \gg n\). The rank-
latent_dimensions SVD of \(\hat{B}\) gives the canonical directions,
which are then whitened so that the canonical variates have unit
variance, sign-aligned to positive correlation, and sorted in
descending order.
Because the penalty acts on rows of \(B\), sparsity is induced only in
\(X\); \(Y\) is handled densely via its inverse-square-root covariance.
Swap the order of views to regularise the other view instead.
This is a NumPy port of the reference R implementation,
ccar3, reusing scikit-learn's
own :class:~sklearn.linear_model.MultiTaskLasso in place of the R
package's CVXR/rrpack solver backends; use GridSearchCV from
cca_zoo.model_selection to select lambda_ as for any other
estimator.
References
Donnat, C., & Tuzhilina, E. (2024). Canonical Correlation Analysis as Reduced Rank Regression in High Dimensions. arXiv:2405.19539.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means before fitting. Default True. |
True
|
lambda_
|
float
|
Row-group-lasso regularisation strength used when
|
0.0
|
highdim
|
bool
|
Whether to estimate the reduced-rank coefficient with the
group-lasso penalty (default, needed when |
True
|
ledoit_wolf
|
bool
|
Whether to shrink the |
True
|
max_iter
|
int
|
Maximum number of coordinate-descent iterations used when
|
10000
|
tol
|
float
|
Convergence tolerance used when |
0.0001
|
eps
|
float
|
Small constant added to covariance matrices before inversion, for numerical stability. Default 1e-8. |
1e-08
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = CCAR3(latent_dimensions=2, highdim=False).fit([X1, X2])
>>> scores = model.transform([X1, X2])
Source code in cca_zoo/linear/_ccar3.py
fit ¶
Fit the CCAR3 model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of exactly two arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
CCAR3
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the number of views is not exactly 2. |
ValueError
|
If views have inconsistent numbers of samples. |
Source code in cca_zoo/linear/_ccar3.py
ECCA ¶
ECCA(
latent_dimensions: int = 1,
center: bool = True,
lambda_: float = 0.0,
max_iter: int = 10000,
tol: float = 0.0001,
eps: float = 1e-08,
)
Bases: BaseModel
Canonical Correlation Analysis via entrywise-sparse Reduced Rank Regression.
Like :class:~cca_zoo.linear.CCAR3, recasts two-view CCA as a reduced
-rank regression of X onto Y, but with an entrywise penalty on
the coefficient matrix \(B\) instead of a row-group one, and -- unlike
CCAR3 -- fit directly against raw (centred) \(Y\), with no
Ledoit-Wolf pre-whitening step:
Where :class:~cca_zoo.linear.CCAR3's row-group-lasso penalty zeroes
whole \(X\) features at once (a feature is either used by every canonical
variate or by none), this entrywise penalty can zero individual
(feature, component) entries independently -- a feature can
contribute to component 1 while being dropped from component 2. This
exactly mirrors the relationship between
:class:~cca_zoo.sparse.ElasticNetCCA (entrywise) and
:class:~cca_zoo.sparse.MultiTaskElasticNetCCA (row-group) one level
up, but here for the reduced-rank-regression family rather than the
Eckart-Young-loss family. The rank-latent_dimensions SVD of
\(\hat{B}\) gives the canonical directions, whitened so the canonical
variates have unit variance, sign-aligned to positive correlation, and
sorted in descending order -- the same postprocessing CCAR3 uses,
just with no Y-covariance un-whitening step (there's no Y-whitening to
undo here).
This is a NumPy port of the reference R implementation's ecca()
function (ccar3), solved by
a bank of independent :class:~sklearn.linear_model.Lasso fits (one
per column of \(Y\)) rather than the R package's single matrix-free ADMM
over the whole \(B\) at once -- the entrywise penalty makes the columns
of \(B\) independent regardless of solver (see
:func:_entrywise_sparse_rrr's docstring), so this reaches the same
optimum with a far simpler, already-well-tested solver. The absence of
Y-whitening (unlike CCAR3) is deliberate, not an oversight: the R
reference's own Sy/LW_Sy machinery is present in cca_rrr()
but explicitly not used by ecca() -- its ecca_across_lambdas
keeps an Sy parameter only "for compatibility" and ignores it. The
R package's optional block/graph groups argument (arbitrary
(x, y) index pairs sharing one penalty) is out of scope here; use
GridSearchCV from cca_zoo.model_selection to select lambda_ as
for any other estimator.
References
ccar3 <https://github.com/jameschapman19/ccar3>_'s ecca(),
the entrywise-sparse companion to
Donnat, C., & Tuzhilina, E. (2024). Canonical Correlation Analysis
as Reduced Rank Regression in High Dimensions. arXiv:2405.19539.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means before fitting. Default True. |
True
|
lambda_
|
float
|
Entrywise lasso regularisation strength. |
0.0
|
max_iter
|
int
|
Maximum number of coordinate-descent iterations, passed
straight through to :class: |
10000
|
tol
|
float
|
Convergence tolerance, passed straight through to
:class: |
0.0001
|
eps
|
float
|
Small constant added to covariance matrices before inversion, for numerical stability. Default 1e-8. |
1e-08
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = ECCA(latent_dimensions=2, lambda_=0.1).fit([X1, X2])
>>> scores = model.transform([X1, X2])
Source code in cca_zoo/linear/_ecca.py
fit ¶
Fit the ECCA model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of exactly two arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
ECCA
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the number of views is not exactly 2. |
ValueError
|
If views have inconsistent numbers of samples. |
Source code in cca_zoo/linear/_ecca.py
Sparse-precision covariance methods¶
GraphicalLassoCCA ¶
GraphicalLassoCCA(
latent_dimensions: int = 1,
center: bool = True,
c: float | list[float] = 0.0,
alpha: float | list[float | None] | None = 0.01,
mode: str = "cd",
max_iter: int = 100,
eps: float = 1e-06,
)
Bases: MCCA
GraphicalLassoCCA -- MCCA with a sparse-precision within-view covariance.
Every within-view regularisation already in this library changes how
the covariance :math:X_i^\top X_i is estimated before it goes into
:class:~cca_zoo.linear.MCCA's generalised eigenproblem -- a fixed
ridge blend toward the identity (MCCA's own c), Ledoit-Wolf
shrinkage (:class:~cca_zoo.linear.CCAR3's ledoit_wolf), or
concentration-step trimming (:class:~cca_zoo.linear.TrimmedCCA).
None of them touch the inverse covariance directly.
GraphicalLassoCCA does: each view's block of
:class:~cca_zoo.linear.MCCA's within-view matrix :math:B is built
from :class:sklearn.covariance.GraphicalLasso's (or, with
alpha=None, :class:sklearn.covariance.GraphicalLassoCV's)
L1-penalised precision estimate's implied covariance, in place of the
raw sample covariance -- an L1 penalty on each view's partial
correlations (conditional independence structure) rather than an L2
shrinkage of the covariance itself. The between-view matrix :math:A
is untouched (plain sample cross-covariance, as in MCCA), so only
the "how confidently does this view's own covariance matrix invert"
side of the eigenproblem changes.
Since the point of estimating a sparse precision matrix is usually the
sparse structure itself, not a dimensionality-reduced approximation of
it, this always solves the eigenproblem directly in each view's
original feature space (:class:~cca_zoo.linear.MCCA's
pca=True shortcut is not applicable here and isn't exposed).
Note
:class:sklearn.covariance.GraphicalLasso (and its CV variant)
estimate a sparse precision matrix under a Gaussian assumption
and are themselves most useful in the high-dimensional
(:math:p \gtrsim n) regime a plain sample covariance can't
invert reliably -- exactly where MCCA's own docs recommend
pca=True instead. This is a different way to make that same
regime tractable: constrain the inverse covariance's structure
rather than truncate the covariance's rank.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means before fitting. Default True. |
True
|
c
|
float | list[float]
|
Ridge blend applied on top of the graphical-lasso covariance
estimate, same semantics as :class: |
0.0
|
alpha
|
float | list[float | None] | None
|
Graphical-lasso L1 penalty strength(s). A scalar or per-view
list of non-negative floats, or |
0.01
|
mode
|
str
|
Graphical-lasso solver, |
'cd'
|
max_iter
|
int
|
Maximum graphical-lasso iterations. Default 100. |
100
|
eps
|
float
|
Small constant added to the eigenvalues of |
1e-06
|
References
Friedman, J., Hastie, T., & Tibshirani, R. (2008). Sparse inverse covariance estimation with the graphical lasso. Biostatistics, 9(3), 432-441.
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((100, 12))
>>> X2 = rng.standard_normal((100, 9))
>>> model = GraphicalLassoCCA(alpha=0.1).fit([X1, X2])
>>> scores = model.transform([X1, X2])
>>> precisions = model.precision_ # sparse per-view precision matrices
Source code in cca_zoo/linear/_graphical_lasso_cca.py
Robust methods¶
RANSACCCA ¶
RANSACCCA(
latent_dimensions: int = 1,
center: bool = True,
c: float | list[float] = 0.1,
min_samples: int | float = 0.25,
residual_threshold: float | None = None,
max_trials: int = 200,
stop_probability: float = 0.99,
random_state: int | None = None,
)
Bases: BaseModel
RANSACCCA -- robust multiview CCA via random sample consensus.
The EY-loss family already has :class:~cca_zoo.linear.HuberCCA for
robustness against high-leverage samples: points whose combined
magnitude across views dominates a covariance-based statistic simply
by being large. That leaves a different failure mode untouched: a
subset of samples whose cross-view relationship is wrong -- mismatched,
corrupted, or drawn from an altogether different relationship -- while
remaining completely ordinary in magnitude within each view on its own,
so nothing about their individual norm flags them as unusual.
:class:~sklearn.linear_model.RANSACRegressor was built for exactly
this in ordinary regression (structured minorities that leverage-based
downweighting can't see), and the same idea carries over here.
Fit follows classical RANSAC: repeatedly draw a random subset of
min_samples rows, fit a candidate model on just that subset with
:class:~cca_zoo.linear.MCCA (a fast closed-form generalised
eigenvalue solve, needed since this repeats many times), and score the
candidate by projecting every sample through it and summing each
sample's positive :func:_cross_view_agreement -- literally, how much
of the data this candidate's direction actually explains. The
best-scoring candidate's inlier set (samples with agreement at or
above residual_threshold) is kept, and the final model is refit on
that consensus set alone. The number of trials adapts to the best
inlier fraction found so far via the standard RANSAC formula, capped
at max_trials.
Note
Like :class:~cca_zoo.linear.MCCA, this is not convex, and here
there is the added instability of the random subset draws
themselves: different random_state seeds can find different
consensus sets, especially when the "wrong" relationship is
supported by close to half the data (see the class's tests for a
worked example of where this method helps and where the problem
itself becomes too ambiguous for any method to resolve reliably).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means. Default is True. |
True
|
c
|
float | list[float]
|
Ridge regularisation passed to every internal
:class: |
0.1
|
min_samples
|
int | float
|
Size of each random candidate subset, as a fraction
of the training set (float in |
0.25
|
residual_threshold
|
float | None
|
Minimum :func: |
None
|
max_trials
|
int
|
Maximum number of random subsets to try. Default is 200. |
200
|
stop_probability
|
float
|
Trials stop early, before |
0.99
|
random_state
|
int | None
|
Seed for the random subset draws. |
None
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((200, 8))
>>> X2 = rng.standard_normal((200, 6))
>>> model = RANSACCCA(latent_dimensions=1, random_state=0).fit([X1, X2])
>>> scores = model.transform([X1, X2])
Source code in cca_zoo/linear/_ransac_cca.py
fit ¶
Fit RANSACCCA by random sample consensus over MCCA candidate fits.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of 2 or more arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
RANSACCCA
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If fewer than 2 views are provided. |
ValueError
|
If views have inconsistent numbers of samples. |
Source code in cca_zoo/linear/_ransac_cca.py
TrimmedCCA ¶
TrimmedCCA(
latent_dimensions: int = 1,
center: bool = True,
c: float = 0.1,
h_frac: float = 0.75,
n_starts: int = 10,
max_iter: int = 30,
tol: float = 1e-08,
random_state: int | None = None,
)
Bases: BaseModel
TrimmedCCA -- robust multiview CCA via concentration steps.
:class:~cca_zoo.linear.RANSACCCA searches for a clean subset by
drawing many small random candidates and keeping the best-scoring
one -- a good strategy while contamination stays well below its
search's own odds of ever drawing a clean-enough sample. As
contamination approaches the ~50% breakdown point, that search
degrades: a random small subset becomes close to a coin flip on
being usably clean, however many trials are tried. TrimmedCCA
instead uses concentration steps in the style of Rousseeuw's Least
Trimmed Squares / Minimum Covariance Determinant: starting from a
large random subset of h rows (h_frac of the data), it
alternates
- select: rank every sample by its own contribution to
:class:
~cca_zoo.linear.gradient.CCAEY's exact loss (for the current weights) and keep the besth-- solved via a Lagrangian relaxation of the loss's own algebraic structure (see :func:_select), not an absolute per-sample threshold likeRANSACCCA's; - refit: re-minimise
CCAEY's exact loss restricted to the kepthrows, warm-started at the current weights via L-BFGS-B.
Each step only ever accepts a subset/weight pair that doesn't
increase the loss (refit is monotone by construction; a selection
step that fails to improve is rejected and the loop stops there) --
the classical C-step argument, applied to CCAEY's real objective
rather than a proxy score. Repeated over n_starts random restarts
(this objective is non-convex, so a single start can land on a poor
local optimum), keeping the lowest-loss result.
Note
A different classical way to get a high-breakdown robust CCA is to
plug the minimum covariance determinant (MCD) estimator into the
joint covariance matrix of both views and solve the ordinary CCA
eigenproblem on that robust estimate (Croux & Dehon, 2002; see also
the comparative study of Branco, Croux, Filzmoser & Oliveira, 2005).
TrimmedCCA differs from that plug-in approach: its concentration
steps minimise CCAEY's own loss directly on the kept subset
(see :func:_select, :func:_refit) rather than the covariance
matrix's determinant as an intermediate, general-purpose target,
so the trimming is targeted at what actually drives the CCA
objective rather than at multivariate location/scatter generally.
h_frac is not learned from the data -- like
:class:sklearn.covariance.MinCovDet's support_fraction, it
is a prior on how much of the training data you expect is
contaminated, set before fitting. Too high wastes some of a
fixed-size budget on good rows discarded unnecessarily when
contamination is actually low; too low forces contaminated rows
into every fit once true contamination exceeds 1 - h_frac.
It cannot be chosen by cross-validating a downstream metric,
since that would need labels for which rows are contaminated --
exactly what's unknown.
TrimmedCCA supports any number of views (2 or more) but only
latent_dimensions=1. The selection rule's closed-form
derivation (see :func:_per_sample_terms) relies on CCAEY's
penalty term being the square of a single linear functional of
the selection -- true regardless of the number of views, but not
past one latent dimension: with \(k > 1\) latent dimensions the
same penalty becomes a genuine matrix-valued quadratic form (rank
up to \(k(k+1)/2\)), which the same single-multiplier bisection
cannot solve. RANSACCCA (via :class:~cca_zoo.linear.MCCA)
supports any number of latent dimensions directly, and matches or
beats TrimmedCCA away from the ~50% breakdown regime -- reach
for TrimmedCCA specifically when contamination is expected to
be heavy and h_frac can be set close to the true clean
fraction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Must be 1 (the only value currently
supported; see the |
1
|
center
|
bool
|
Whether to subtract column means. Default True. |
True
|
c
|
float
|
Ridge blend in |
0.1
|
h_frac
|
float
|
Fraction of rows kept every concentration step, in
|
0.75
|
n_starts
|
int
|
Random restarts; the lowest-loss result is kept. Default 10. |
10
|
max_iter
|
int
|
Maximum concentration steps per restart. Default 30. |
30
|
tol
|
float
|
Convergence tolerance for each refit's L-BFGS-B call, passed
as |
1e-08
|
random_state
|
int | None
|
Seed for the random restarts. |
None
|
References
Rousseeuw, P. J., & Van Driessen, K. (1999). A fast algorithm for the minimum covariance determinant estimator. Technometrics, 41(3), 212-223.
Croux, C., & Dehon, C. (2002). Analyse canonique basee sur des estimateurs robustes de la matrice de covariance. Revue de Statistique Appliquee, 50(2), 5-26.
Branco, J. A., Croux, C., Filzmoser, P., & Oliveira, M. R. (2005). Robust canonical correlations: A comparative study. Computational Statistics, 20(2), 203-229.
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((200, 8))
>>> X2 = rng.standard_normal((200, 6))
>>> model = TrimmedCCA(h_frac=0.7, random_state=0).fit([X1, X2])
>>> inliers = model.inlier_mask_ # boolean array over the training rows
More than two views are supported directly:
>>> X3 = rng.standard_normal((200, 5))
>>> model = TrimmedCCA(h_frac=0.7, random_state=0).fit([X1, X2, X3])
Source code in cca_zoo/linear/_trimmed_cca.py
fit ¶
Fit TrimmedCCA by concentration steps on CCAEY's exact loss.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of 2 or more arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
TrimmedCCA
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If fewer than 2 views are provided. |
ValueError
|
If |
Source code in cca_zoo/linear/_trimmed_cca.py
ProjectionPursuitCCA ¶
ProjectionPursuitCCA(
latent_dimensions: int = 1,
center: bool = True,
projection_index: str = "spearman",
mcd_support_fraction: float = 0.75,
n_restarts: int = 10,
max_iter: int = 200,
tol: float = 1e-06,
random_state: int | None = None,
)
Bases: BaseModel
ProjectionPursuitCCA -- robust multiview CCA via projection pursuit.
Every other estimator in :mod:cca_zoo.linear -- robust or not -- is
built from a cross- or auto-covariance statistic of the projected
views, computed once the projection directions are (implicitly or
explicitly) fixed. Projection pursuit inverts that: it never forms a
covariance matrix at all, instead searching directly over candidate
projection directions for the pair that maximises a robust bivariate
correlation measure (the projection index) between the resulting
univariate scores, following the classical projection-pursuit paradigm
of Huber (1985) as carried over to CCA specifically by Branco, Croux,
Filzmoser & Oliveira (2005) and put on firmer statistical footing
(efficiency, breakdown point, and a wider family of projection indices)
by Alfons, Croux & Filzmoser (2017). Reference implementation: the R
package ccaPP <https://cran.r-project.org/package=ccaPP>_ (Alfons,
Croux & Filzmoser, 2016, Austrian Journal of Statistics).
For two views this reduces to their exact problem: find unit vectors
\(\mathbf{a}, \mathbf{b}\) maximising \(\operatorname{PI}(X\mathbf{a},
Y\mathbf{b})\) for a robust correlation measure \(\operatorname{PI}\).
ProjectionPursuitCCA generalises this to \(M \geq 2\) views by
maximising the average projection index over every pair of views
(the same generalisation :class:~cca_zoo.linear.RANSACCCA's
consensus score and :class:~cca_zoo.linear.MCCA's sum-of-pairwise
objective both make):
Two projection indices are available: "spearman" (default, Spearman
rank correlation -- see :func:spearman_projection_index) and "mcd"
(a minimum-covariance-determinant-based correlation -- see
:func:mcd_projection_index).
Fit by direct numerical search: each unit vector is parametrised by
\(p_i - 1\) unconstrained angles (:func:_angles_to_unit_vector), and
the (generally non-smooth -- a rank correlation changes discontinuously
wherever two projected scores swap rank order) objective above is
maximised over the stacked angle vector by Powell's method
(derivative-free, unlike the L-BFGS-B used elsewhere in this package,
since the objective's gradient is undefined almost everywhere) from
n_restarts random starting points, keeping the best. Later latent
dimensions are fit the same way on views deflated by the previously
found directions (:func:~cca_zoo._utils._linalg.deflate, the same
Gram-Schmidt convention :mod:cca_zoo.sparse's ALS-based methods use).
Note
Unlike :class:~cca_zoo.linear.HuberCCA (leverage-based) and
:class:~cca_zoo.linear.RANSACCCA/:class:~cca_zoo.linear.TrimmedCCA
(both relational, via a per-sample loss or agreement score),
projection pursuit does not single out individual bad rows at all
-- its robustness comes entirely from the projection index itself
being insensitive to a handful of extreme values, whatever kind of
contamination produced them. There is accordingly no
inlier_mask_ to inspect after fitting.
The random-restart search is not guaranteed to find the global
optimum (the objective is non-convex and, for "spearman",
genuinely discontinuous), and its cost scales with n_restarts
times the cost of one projection-index evaluation times
latent_dimensions; "mcd" is markedly more expensive per
evaluation than "spearman" since it re-fits a robust covariance
estimator at every candidate direction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means. Default is True. |
True
|
projection_index
|
str
|
|
'spearman'
|
mcd_support_fraction
|
float
|
Passed to :class: |
0.75
|
n_restarts
|
int
|
Random restarts of the direction search per latent dimension; the best-scoring result is kept. Default 10. |
10
|
max_iter
|
int
|
Maximum Powell iterations per restart. Default 200. |
200
|
tol
|
float
|
Convergence tolerance for Powell's method ( |
1e-06
|
random_state
|
int | None
|
Seed for the random restarts (and for |
None
|
References
Huber, P. J. (1985). Projection pursuit. The Annals of Statistics, 13(2), 435-475.
Branco, J. A., Croux, C., Filzmoser, P., & Oliveira, M. R. (2005). Robust canonical correlations: A comparative study. Computational Statistics, 20(2), 203-229.
Alfons, A., Croux, C., & Filzmoser, P. (2017). Robust maximum association estimators. Journal of the American Statistical Association, 112(517), 436-445.
Alfons, A., Croux, C., & Filzmoser, P. (2016). Robust maximum association between data sets: The R package ccaPP. Austrian Journal of Statistics, 45(1), 71-79.
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((200, 6))
>>> X2 = rng.standard_normal((200, 5))
>>> model = ProjectionPursuitCCA(latent_dimensions=1, random_state=0)
>>> model = model.fit([X1, X2])
Source code in cca_zoo/linear/_projection_pursuit_cca.py
fit ¶
Fit ProjectionPursuitCCA by direct search over projection directions.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of 2 or more arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
ProjectionPursuitCCA
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If fewer than 2 views are provided. |
ValueError
|
If views have inconsistent numbers of samples. |
Source code in cca_zoo/linear/_projection_pursuit_cca.py
EY-loss methods¶
PLSEY ¶
PLSEY(
latent_dimensions: int = 1,
center: bool = True,
max_iter: int = 1000,
tol: float = 1e-08,
random_state: int | None = None,
)
Bases: CCAEY
Eckart-Young PLS for 2 or more views.
This is equivalent to :class:~cca_zoo.linear.gradient.CCAEY with
c=1: the reward excludes the \(i = j\) terms that CCAEY's
(\(c=0\)) reward includes, and the penalty is purely
\(\operatorname{tr}(BB)\) on the weight Gram matrix \(B\), which
drives the weights towards (approximate) orthonormality at the optimum
on its own — no manifold projection step, and no upfront whitening.
Suitable for high-dimensional data where forming the full (\(p \times p\))
cross-covariance matrix is too expensive. Fit by full-batch L-BFGS-B
using the loss's exact analytic gradient; for mini-batch training on
datasets too large for a full-batch gradient evaluation, see
:class:~cca_zoo.linear.gradient.StochasticCCAEY (c=1).
Initial weights have exactly orthonormal columns (unit-norm, mutually
orthogonal) before any optimisation step, matching the shape of this
loss's own penalty on \(B\) — unlike :class:~cca_zoo.linear.gradient.CCAEY's
own data-informed default, which instead orthonormalises the initial
projections (see :func:cca_zoo._utils._ey.random_orthonormal_weights
vs. :func:cca_zoo._utils._ey.cheap_orthonormal_projection_weights).
References
Chapman, J., Wells, L., & Lawry Aguila, A. (2024). Unconstrained Stochastic CCA: Unifying Multiview and Self-Supervised Learning. arXiv:2310.01012.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means. Default True. |
True
|
max_iter
|
int
|
Maximum number of L-BFGS-B iterations. Default is 1000. |
1000
|
tol
|
float
|
Convergence tolerance, passed to L-BFGS-B as |
1e-08
|
random_state
|
int | None
|
Seed for reproducibility. |
None
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((200, 500))
>>> X2 = rng.standard_normal((200, 400))
>>> model = PLSEY(latent_dimensions=4, random_state=0)
>>> model = model.fit([X1, X2])
Source code in cca_zoo/linear/gradient/_pls_ey.py
fit ¶
Fit PLSEY by full-batch L-BFGS-B on the EY loss.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of 2 or more arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
PLSEY
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If fewer than 2 views are provided. |
ValueError
|
If views have inconsistent numbers of samples. |
Source code in cca_zoo/linear/gradient/_pls_ey.py
CCAEY ¶
CCAEY(
latent_dimensions: int = 1,
center: bool = True,
c: float = 0.0,
max_iter: int = 1000,
tol: float = 1e-08,
random_state: int | None = None,
)
Bases: BaseFullBatchEYModel
Eckart-Young CCA for 2 or more views, ridge-blended with PLSEY.
Minimises the unconstrained Eckart-Young (EY) objective directly on the
raw (centred) views, with no manifold projection step and no upfront
whitening: unlike classical CCA, which whitens each view before finding
the correlated directions, the EY reformulation folds the
orthonormalising pressure into the loss itself, so a full-batch
preprocessing pass over the data covariance is never needed. This
matches how the same underlying loss is used, unwhitened, by
:class:~cca_zoo.linear.gradient.PLSEY, :class:~cca_zoo.tree.TreeCCA,
and :class:~cca_zoo.deep.DCCAEY.
For embeddings \(Z_i = X_i W_i\) (\(i = 1, \dots, M\), \(M \ge 2\)), let \(C\)
and \(V\) be the mean pairwise cross-covariance and mean auto-covariance
across views (see :func:cca_zoo._utils._ey.ey_cross_covariance), and
\(B = \frac{1}{M}\sum_i W_i^\top W_i\) the mean weight Gram matrix
(see :func:cca_zoo._utils._ey.weight_gram_mean). c blends the
within-view normalisation between the data's own auto-covariance and
the identity (in weight space, \(W_i^\top I W_i = W_i^\top W_i\)) —
exactly the canonical-ridge blend \((1-c)X^\top X + cI\) already used by
:class:~cca_zoo.linear.rCCA, translated into this unconstrained
setting:
c=0 recovers plain (unregularised) CCAEY exactly; c=1
recovers :class:~cca_zoo.linear.gradient.PLSEY's loss exactly (its
reward excludes the \(i=j\) terms that \(\mathcal{L}_{EY}(0)\)
includes, and its penalty is purely \(\operatorname{tr}(BB)\)) —
both endpoints, and the gradient at intermediate \(c\), are verified
against finite differences and against PLSEY's own independently
verified gradient. This objective has the canonical directions as a
stationary point without requiring an explicit orthonormality
constraint, unlike a plain squared-projection-distance loss.
Fit by full-batch L-BFGS-B
(:meth:~cca_zoo.linear.gradient._base.BaseFullBatchEYModel._fit_lbfgsb)
using the loss's exact analytic gradient. For mini-batch training on
datasets too large for a full-batch gradient evaluation, see
:class:~cca_zoo.linear.gradient.StochasticCCAEY.
Note
Unlike the exact, closed-form :class:~cca_zoo.linear.rCCA (where
c=0 is always numerically safe), optimising the raw,
unregularised (\(c=0\)) objective can be poorly conditioned when the
number of samples doesn't outnumber the number of features by a
healthy margin, since nothing then bounds the weights in the data's
near-null directions. If you see nan or diverging weights,
increase c (a small value like 0.1-0.3 is usually enough).
References
Chapman, J., Wells, L., & Lawry Aguila, A. (2024). Unconstrained Stochastic CCA: Unifying Multiview and Self-Supervised Learning. arXiv:2310.01012.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means. Default True. |
True
|
c
|
float
|
Ridge blend in |
0.0
|
max_iter
|
int
|
Maximum number of L-BFGS-B iterations. Default is 1000. |
1000
|
tol
|
float
|
Convergence tolerance, passed to L-BFGS-B as |
1e-08
|
random_state
|
int | None
|
Seed for reproducibility. |
None
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((1000, 20))
>>> X2 = rng.standard_normal((1000, 15))
>>> model = CCAEY(latent_dimensions=4, random_state=0)
>>> model = model.fit([X1, X2])
More than two views are supported directly:
>>> X3 = rng.standard_normal((1000, 10))
>>> model = CCAEY(latent_dimensions=4, random_state=0).fit([X1, X2, X3])
Source code in cca_zoo/linear/gradient/_cca_ey.py
fit ¶
Fit CCAEY by full-batch L-BFGS-B on the EY loss.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of 2 or more arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
CCAEY
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If fewer than 2 views are provided. |
ValueError
|
If views have inconsistent numbers of samples. |
Source code in cca_zoo/linear/gradient/_cca_ey.py
HuberCCA ¶
HuberCCA(
latent_dimensions: int = 1,
center: bool = True,
delta: float = 4.0,
max_iter: int = 1000,
tol: float = 1e-08,
random_state: int | None = None,
)
Bases: BaseFullBatchEYModel
Huber CCA: bounded-influence Eckart-Young CCA.
Standard :class:~cca_zoo.linear.gradient.CCAEY weights every sample
equally, so its cross- and auto-covariance estimates (and hence its
gradient) can be dominated by a handful of high-leverage points -- their
contribution to a quadratic statistic grows with the square of their
magnitude, unbounded. This is the same weakness ordinary least squares
has relative to :class:~sklearn.linear_model.HuberRegressor: quadratic
loss growth lets outliers dominate, so Huber loss caps it to linear
growth instead.
HuberCCA reweights each sample by a Huber-style factor of its own
leverage (see :func:_huber_sample_weight) before forming the EY
cross- and auto-covariance statistics (see :func:_weighted_ey):
samples within delta times the dataset's own median leverage keep
weight 1, samples beyond it are downweighted so their contribution is
capped rather than unbounded -- every sample still contributes something,
just never an unbounded amount. Fit by full-batch L-BFGS-B
(:meth:~cca_zoo.linear.gradient._base.BaseFullBatchEYModel._fit_lbfgsb),
exactly as :class:~cca_zoo.linear.gradient.CCAEY; the leverage weight
is recomputed from the current representations at every evaluation,
the same stop-gradient (IRLS-style) treatment described in
:func:_weighted_ey.
Note
Like plain CCAEY at its unregularised c=0 (this estimator
has no ridge-blend c of its own), optimising this objective can
be poorly conditioned when the number of samples doesn't outnumber
the number of features by a healthy margin. If you see nan or
diverging weights, this is the same caveat CCAEY documents, not
specific to the Huber reweighting.
Reweighting a CCA fit by each sample's own leverage is not a new
idea: Filzmoser, Dehon & Croux (2000) and Branco, Croux, Filzmoser &
Oliveira (2005) already reweight Wold's (1966) alternating-regression
CCA fit by robust-distance-based weights on each side in turn
("Robust Alternating Regression"), and Croux & Dehon (2002) instead
plug a minimum covariance determinant estimator directly into the
classical covariance-matrix eigenproblem. HuberCCA's difference
is mechanism, not the underlying bounded-influence idea: it
reweights the unconstrained EY loss's own sample statistics
directly (see :func:_huber_sample_weight, :func:_weighted_ey)
rather than an alternating regression or a general-purpose
covariance estimator, with a cutoff self-calibrated to the batch's
own median leverage rather than an MVE- or M-scale-based threshold.
References
Filzmoser, P., Dehon, C., & Croux, C. (2000). Outlier resistant estimators for canonical correlation analysis. In COMPSTAT: Proceedings in Computational Statistics 2000 (pp. 301-306). Physica-Verlag.
Branco, J. A., Croux, C., Filzmoser, P., & Oliveira, M. R. (2005). Robust canonical correlations: A comparative study. Computational Statistics, 20(2), 203-229.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent_dimensions
|
int
|
Number of latent dimensions. Default is 1. |
1
|
center
|
bool
|
Whether to subtract column means. Default True. |
True
|
delta
|
float
|
Huber cutoff, as a multiple of the dataset's median sample leverage; samples beyond it are downweighted. Values below 1 downweight the majority of the data and are not recommended. Smaller values are more robust but discard more of the data's genuine signal. Default is 4.0. |
4.0
|
max_iter
|
int
|
Maximum number of L-BFGS-B iterations. Default is 1000. |
1000
|
tol
|
float
|
Convergence tolerance, passed to L-BFGS-B as |
1e-08
|
random_state
|
int | None
|
Seed for reproducibility. |
None
|
Examples:
>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((1000, 20))
>>> X2 = rng.standard_normal((1000, 15))
>>> model = HuberCCA(latent_dimensions=4, random_state=0)
>>> model = model.fit([X1, X2])
Source code in cca_zoo/linear/gradient/_huber_cca.py
fit ¶
Fit HuberCCA by full-batch L-BFGS-B on the Huber-weighted EY loss.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
views
|
list[ArrayLike]
|
List of arrays, each (n_samples, n_features_i). |
required |
y
|
None
|
Ignored. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
HuberCCA
|
Fitted estimator. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If fewer than 2 views are provided. |
ValueError
|
If views have inconsistent numbers of samples. |