Skip to content

cca_zoo.sparse

Sparse linear CCA methods. No optional dependency required. Two mechanism families live here: EY-loss coordinate descent, and Alternating Least Squares (ALS).


EY-loss coordinate descent

ElasticNetCCA

ElasticNetCCA(
    latent_dimensions: int = 1,
    center: bool = True,
    alpha: float | list[float] = 1.0,
    l1_ratio: float | list[float] = 0.5,
    max_iter: int = 100,
    tol: float = 1e-06,
    random_state: int | None = None,
    positive: bool = False,
)

Bases: BaseModel

ElasticNetCCA — sparse linear CCA by coordinate descent on the EY loss.

Learns per-view linear weights \(W_i\) (embeddings \(Z_i = X_i W_i\)) that minimise the elastic-net-penalised Eckart-Young (EY) objective:

\[ \mathcal{L}(W) = \mathcal{L}_{EY}(Z_1, \dots, Z_M) + \sum_i \left( \alpha \, \rho \, \|W_i\|_1 + \tfrac{1}{2} \alpha (1-\rho) \|W_i\|_F^2 \right) \]

where \(\mathcal{L}_{EY}\) is the EY loss (see :mod:cca_zoo._utils._ey, shared with :class:~cca_zoo.linear.gradient.CCAEY, :class:~cca_zoo.tree.TreeCCA, :class:~cca_zoo.gam.GAMCCA, and :class:~cca_zoo.gp.GaussianProcessCCA) and \(\rho\) is l1_ratio. This is fit by :func:~cca_zoo._utils._ey.coordinate_descent_ey — cyclic coordinate descent on the EY loss itself (bases = the raw centred views, no ridge coupling), the same algorithm :class:~sklearn.linear_model.ElasticNet uses for ordinary (squared-error) elastic net, but with each coordinate's exact minimiser solved against \(\mathcal{L}_{EY}\)'s own (quartic, not quadratic) restriction — see that function's docstring for the derivation.

Each embedding \(Z_i\) is exactly linear in \(X_i\) throughout fitting, so it inherits :class:~cca_zoo._base.BaseModel's plain transform/weights machinery unmodified, and weights genuinely are the sparse canonical weight vectors — not a placeholder that raises NotImplementedError the way it does for :class:~cca_zoo.tree.TreeCCA.

Note

Like every EY-loss model, \(\mathcal{L}_{EY}\) is not convex in \(W\) jointly (only each single coordinate's restriction is, in the limited sense of being an exactly-solvable quartic), so coordinate descent is only guaranteed to reach a stationary point, and different random_state initialisations can land on different ones — the same caveat that already applies to :class:~cca_zoo.linear.gradient.CCAEY's gradient descent.

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent dimensions. Default is 1.

1
center bool

Whether to subtract column means. Default is True.

True
alpha float | list[float]

Overall elastic-net penalty strength(s). Either a single float applied to every view or a list of per-view floats. Default is 1.0.

1.0
l1_ratio float | list[float]

Elastic-net mixing parameter(s) in [0, 1]; 0 is pure ridge, 1 is pure lasso. Either a single float applied to every view or a list of per-view floats. Default is 0.5.

0.5
max_iter int

Maximum number of full coordinate-descent sweeps (every view, feature, and component once each). Default is 100.

100
tol float

Convergence tolerance on the penalised objective's change between consecutive sweeps. Default is 1e-6.

1e-06
random_state int | None

Seed for the initial weights.

None
positive bool

If True, constrain every weight to be non-negative (mirrors :class:~sklearn.linear_model.Lasso and :class:~sklearn.linear_model.ElasticNet's positive=True). Default is False.

False

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((200, 20))
>>> X2 = rng.standard_normal((200, 15))
>>> model = ElasticNetCCA(latent_dimensions=2, alpha=0.1).fit([X1, X2])
>>> scores = model.transform([X1, X2])

A different penalty per view:

>>> model = ElasticNetCCA(latent_dimensions=2, alpha=[0.1, 0.5]).fit(
...     [X1, X2]
... )
Source code in cca_zoo/sparse/_elasticnetcca.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    alpha: float | list[float] = 1.0,
    l1_ratio: float | list[float] = 0.5,
    max_iter: int = 100,
    tol: float = 1e-6,
    random_state: int | None = None,
    positive: bool = False,
) -> None:
    super().__init__(latent_dimensions=latent_dimensions, center=center)
    self.alpha = alpha
    self.l1_ratio = l1_ratio
    self.max_iter = max_iter
    self.tol = tol
    self.random_state = random_state
    self.positive = positive

fit

fit(
    views: list[ArrayLike], y: None = None
) -> ElasticNetCCA

Fit ElasticNetCCA by cyclic coordinate descent on the EY loss.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of 2 or more arrays, each (n_samples, n_features_i).

required
y None

Ignored.

None

Returns:

Name Type Description
self ElasticNetCCA

Fitted estimator.

Raises:

Type Description
ValueError

If fewer than 2 views are provided.

ValueError

If views have inconsistent numbers of samples.

Source code in cca_zoo/sparse/_elasticnetcca.py
def fit(self, views: list[ArrayLike], y: None = None) -> ElasticNetCCA:
    """Fit ElasticNetCCA by cyclic coordinate descent on the EY loss.

    Args:
        views: List of 2 or more arrays, each (n_samples, n_features_i).
        y: Ignored.

    Returns:
        self: Fitted estimator.

    Raises:
        ValueError: If fewer than 2 views are provided.
        ValueError: If views have inconsistent numbers of samples.
    """
    views_ = self._setup_fit(views)
    alpha_ = perview_parameter("alpha", self.alpha, 1.0, self.n_views_)
    l1_ratio_ = perview_parameter("l1_ratio", self.l1_ratio, 0.5, self.n_views_)
    rng = np.random.default_rng(self.random_state)
    weights, _ = coordinate_descent_ey(
        bases=views_,
        k=self.latent_dimensions,
        alpha=alpha_,
        l1_ratio=l1_ratio_,
        max_iter=self.max_iter,
        tol=self.tol,
        rng=rng,
        positive=self.positive,
    )
    self.weights_: list[np.ndarray] = weights
    return self

MultiTaskElasticNetCCA

MultiTaskElasticNetCCA(
    latent_dimensions: int = 1,
    center: bool = True,
    alpha: float | list[float] = 1.0,
    l1_ratio: float | list[float] = 0.5,
    max_iter: int = 100,
    tol: float = 1e-06,
    random_state: int | None = None,
)

Bases: BaseModel

MultiTaskElasticNetCCA — row-group-sparse linear CCA, fit on the EY loss.

Like :class:~cca_zoo.sparse.ElasticNetCCA, learns per-view linear weights \(W_i\) minimising an elastic-net-penalised EY loss, but with sklearn's :class:~sklearn.linear_model.MultiTaskLasso / :class:~sklearn.linear_model.MultiTaskElasticNet row-group penalty in place of a plain per-scalar penalty:

\[ \mathcal{L}(W) = \mathcal{L}_{EY}(Z_1, \dots, Z_M) + \sum_i \left( \alpha \, \rho \, \|W_i\|_{2,1} + \tfrac{1}{2} \alpha (1-\rho) \|W_i\|_F^2 \right) \]

where \(\|W_i\|_{2,1} = \sum_j \|W_i[j,:]\|_2\) sums each feature's weight-row Euclidean norm over all latent_dimensions components. Because the penalty on a row is zero only when the whole row is zero, a feature is either used by every canonical variate or by none — unlike :class:~cca_zoo.sparse.ElasticNetCCA, which can (and often does) keep a feature for component 1 while dropping it from component 2. That joint selection is exactly what :class:~sklearn.linear_model.MultiTaskLasso buys over plain :class:~sklearn.linear_model.Lasso for ordinary multi-output regression, and it is arguably an even more natural fit here, since a CCA model's latent_dimensions are not independent "tasks" to be fit separately but different views of the same underlying features.

Fit by :func:~cca_zoo._utils._ey.group_coordinate_descent_ey — block-coordinate descent, one row (feature, across all components) at a time, using proximal gradient (ISTA) with backtracking line search rather than :class:~cca_zoo.sparse.ElasticNetCCA's exact per-scalar quartic solve. The two aren't interchangeable: a row's smooth EY restriction couples all \(k\) of its entries together (through the auto-covariance's off-diagonal terms), so there is no closed form for the exact row minimiser the way there is for a single scalar — see that function's docstring for the full derivation. Every accepted step is still a verified decrease of the exact penalised objective, so fitting remains provably monotonic, just without per-step exactness.

Note

Like every EY-loss model, this is not convex in \(W\), so different random_state initialisations can land on different stationary points — see :class:~cca_zoo.sparse.ElasticNetCCA's docstring.

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent dimensions. Default is 1.

1
center bool

Whether to subtract column means. Default is True.

True
alpha float | list[float]

Overall penalty strength(s). Either a single float applied to every view or a list of per-view floats. Default is 1.0.

1.0
l1_ratio float | list[float]

Mixing parameter(s) in [0, 1]; 0 is pure (Frobenius) ridge, 1 is pure row-group lasso. Either a single float applied to every view or a list of per-view floats. Default is 0.5.

0.5
max_iter int

Maximum number of full coordinate-descent sweeps. Default is 100.

100
tol float

Convergence tolerance on the penalised objective's change between consecutive sweeps. Default is 1e-6.

1e-06
random_state int | None

Seed for the initial weights.

None

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((200, 20))
>>> X2 = rng.standard_normal((200, 15))
>>> model = MultiTaskElasticNetCCA(latent_dimensions=2, alpha=0.1).fit([X1, X2])
>>> scores = model.transform([X1, X2])

A different penalty per view:

>>> model = MultiTaskElasticNetCCA(
...     latent_dimensions=2, alpha=[0.1, 0.5]
... ).fit([X1, X2])
Source code in cca_zoo/sparse/_multitaskelasticnetcca.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    alpha: float | list[float] = 1.0,
    l1_ratio: float | list[float] = 0.5,
    max_iter: int = 100,
    tol: float = 1e-6,
    random_state: int | None = None,
) -> None:
    super().__init__(latent_dimensions=latent_dimensions, center=center)
    self.alpha = alpha
    self.l1_ratio = l1_ratio
    self.max_iter = max_iter
    self.tol = tol
    self.random_state = random_state

fit

fit(
    views: list[ArrayLike], y: None = None
) -> MultiTaskElasticNetCCA

Fit MultiTaskElasticNetCCA by row-group coordinate descent on the EY loss.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of 2 or more arrays, each (n_samples, n_features_i).

required
y None

Ignored.

None

Returns:

Name Type Description
self MultiTaskElasticNetCCA

Fitted estimator.

Raises:

Type Description
ValueError

If fewer than 2 views are provided.

ValueError

If views have inconsistent numbers of samples.

Source code in cca_zoo/sparse/_multitaskelasticnetcca.py
def fit(self, views: list[ArrayLike], y: None = None) -> MultiTaskElasticNetCCA:
    """Fit MultiTaskElasticNetCCA by row-group coordinate descent on the EY loss.

    Args:
        views: List of 2 or more arrays, each (n_samples, n_features_i).
        y: Ignored.

    Returns:
        self: Fitted estimator.

    Raises:
        ValueError: If fewer than 2 views are provided.
        ValueError: If views have inconsistent numbers of samples.
    """
    views_ = self._setup_fit(views)
    alpha_ = perview_parameter("alpha", self.alpha, 1.0, self.n_views_)
    l1_ratio_ = perview_parameter("l1_ratio", self.l1_ratio, 0.5, self.n_views_)
    rng = np.random.default_rng(self.random_state)
    weights, _ = group_coordinate_descent_ey(
        bases=views_,
        k=self.latent_dimensions,
        alpha=alpha_,
        l1_ratio=l1_ratio_,
        max_iter=self.max_iter,
        tol=self.tol,
        rng=rng,
    )
    self.weights_: list[np.ndarray] = weights
    return self

OrthogonalMatchingPursuitCCA

OrthogonalMatchingPursuitCCA(
    latent_dimensions: int = 1,
    center: bool = True,
    n_nonzero_coefs: int | list[int] | None = None,
    max_iter: int = 10,
    tol: float = 1e-06,
    random_state: int | None = None,
)

Bases: BaseModel

OrthogonalMatchingPursuitCCA — fixed-cardinality CCA via greedy selection.

Learns per-view linear weights \(W_i\) with each view's number of active (nonzero-row) features capped at a fixed budget, the EY-loss analogue of sklearn's :class:~sklearn.linear_model.OrthogonalMatchingPursuit: where :class:~cca_zoo.sparse.ElasticNetCCA / :class:~cca_zoo.sparse.MultiTaskElasticNetCCA reach a sparsity level indirectly by tuning a continuous penalty strength, this fixes the sparsity level directly by specifying how many features each view may use.

Fit by :func:~cca_zoo._utils._ey.omp_coordinate_descent_ey: for each view, features are added one at a time to a growing active set — each one chosen by the same criterion classical OMP uses (largest-magnitude gradient of the loss at a coefficient of zero, i.e. residual correlation, generalised here from a scalar to the norm of a \(k\)-vector since every feature has one coefficient per latent dimension) — and the active coefficients are re-solved to their exact joint optimum after every addition via :class:~cca_zoo.sparse.ElasticNetCCA's own exact quartic coordinate solve with no penalty. Every feature outside a view's active set is exactly zero; nothing else is shrunk.

Note

Like every EY-loss model, this is not convex, so different random_state initialisations (used only for the warm-started bootstrap the greedy search regrows from — see :func:~cca_zoo._utils._ey.omp_coordinate_descent_ey's docstring for why a literal zero start doesn't work here) can select different active sets.

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent dimensions. Default is 1.

1
center bool

Whether to subtract column means. Default is True.

True
n_nonzero_coefs int | list[int] | None

Target number of active features per view. An int applies the same budget to every view; a list gives one budget per view (must match the number of views passed to fit); None defaults each view to max(1, n_features_i // 10), mirroring sklearn's :class:~sklearn.linear_model.OrthogonalMatchingPursuit default. A budget larger than a view's own feature count is silently capped to that count. Default is None.

None
max_iter int

Maximum number of outer rounds cycling through every view and regrowing its active set from scratch. Default is 10.

10
tol float

Convergence tolerance on the (unpenalised) EY objective's change, both between outer rounds and between the coordinate-descent sweeps used to refit each active set. Default is 1e-6.

1e-06
random_state int | None

Seed for the initial (dense, pre-selection) weights.

None

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((200, 20))
>>> X2 = rng.standard_normal((200, 15))
>>> model = OrthogonalMatchingPursuitCCA(
...     latent_dimensions=2, n_nonzero_coefs=5
... ).fit([X1, X2])
>>> scores = model.transform([X1, X2])
Source code in cca_zoo/sparse/_ompcca.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    n_nonzero_coefs: int | list[int] | None = None,
    max_iter: int = 10,
    tol: float = 1e-6,
    random_state: int | None = None,
) -> None:
    super().__init__(latent_dimensions=latent_dimensions, center=center)
    self.n_nonzero_coefs = n_nonzero_coefs
    self.max_iter = max_iter
    self.tol = tol
    self.random_state = random_state

fit

fit(
    views: list[ArrayLike], y: None = None
) -> OrthogonalMatchingPursuitCCA

Fit OrthogonalMatchingPursuitCCA by greedy forward selection on the EY loss.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of 2 or more arrays, each (n_samples, n_features_i).

required
y None

Ignored.

None

Returns:

Name Type Description
self OrthogonalMatchingPursuitCCA

Fitted estimator.

Raises:

Type Description
ValueError

If fewer than 2 views are provided.

ValueError

If views have inconsistent numbers of samples.

ValueError

If n_nonzero_coefs is a list of the wrong length, or resolves to a non-positive budget for some view.

Source code in cca_zoo/sparse/_ompcca.py
def fit(
    self, views: list[ArrayLike], y: None = None
) -> OrthogonalMatchingPursuitCCA:
    """Fit OrthogonalMatchingPursuitCCA by greedy forward selection on the EY loss.

    Args:
        views: List of 2 or more arrays, each (n_samples, n_features_i).
        y: Ignored.

    Returns:
        self: Fitted estimator.

    Raises:
        ValueError: If fewer than 2 views are provided.
        ValueError: If views have inconsistent numbers of samples.
        ValueError: If ``n_nonzero_coefs`` is a list of the wrong
            length, or resolves to a non-positive budget for some view.
    """
    views_ = self._setup_fit(views)
    n_nonzero_coefs = self._resolve_n_nonzero_coefs(self.n_features_in_)
    rng = np.random.default_rng(self.random_state)
    weights, _ = omp_coordinate_descent_ey(
        bases=views_,
        k=self.latent_dimensions,
        n_nonzero_coefs=n_nonzero_coefs,
        max_iter=self.max_iter,
        tol=self.tol,
        rng=rng,
    )
    self.weights_: list[np.ndarray] = weights
    return self

Alternating Least Squares

PMDCCA

PMDCCA(
    latent_dimensions: int = 1,
    center: bool = True,
    tau: float | list[float] = 1.0,
    max_iter: int = 500,
    tol: float = 1e-06,
    random_state: int | None = None,
)

Bases: _BaseIterative

Sparse CCA via Penalized Matrix Decomposition.

Maximises the cross-view covariance subject to L1 norm constraints on each weight vector:

\[ \begin{aligned} \max_{\mathbf{w}_1, \mathbf{w}_2} \mathbf{w}_1^\top X_1^\top X_2 \mathbf{w}_2 \\ \text{subject to } \|\mathbf{w}_i\|_1 \leq \tau_i \sqrt{p_i},\quad \|\mathbf{w}_i\|_2 = 1 \end{aligned} \]

The update for each view uses bisection to find the soft-threshold that satisfies the L1 constraint exactly.

References

Witten, D. M., Tibshirani, R., & Hastie, T. (2009). A penalized matrix decomposition, with applications to sparse principal components and canonical correlation analysis. Biostatistics, 10(3), 515–534.

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent dimensions. Default is 1.

1
center bool

Whether to subtract column means. Default True.

True
tau float | list[float]

L1 bound scaling factor(s) in (0, 1]. The actual L1 bound is tau * sqrt(n_features_i). Default is 1 (no sparsity).

1.0
max_iter int

Maximum ALS iterations. Default is 500.

500
tol float

Convergence tolerance. Default is 1e-6.

1e-06
random_state int | None

Seed for reproducibility.

None

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = PMDCCA(tau=0.5, random_state=0).fit([X1, X2])
Source code in cca_zoo/sparse/_iterative.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    tau: float | list[float] = 1.0,
    max_iter: int = 500,
    tol: float = 1e-6,
    random_state: int | None = None,
) -> None:
    super().__init__(
        latent_dimensions=latent_dimensions,
        center=center,
        max_iter=max_iter,
        tol=tol,
        random_state=random_state,
    )
    self.tau = tau

fit

fit(views: list[ArrayLike], y: None = None) -> PMDCCA

Fit the PMDCCA model.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of arrays, each (n_samples, n_features_i).

required
y None

Ignored.

None

Returns:

Name Type Description
self PMDCCA

Fitted estimator.

Raises:

Type Description
ValueError

If fewer than 2 views are provided.

Source code in cca_zoo/sparse/_iterative.py
def fit(self, views: list[ArrayLike], y: None = None) -> PMDCCA:
    """Fit the PMDCCA model.

    Args:
        views: List of arrays, each (n_samples, n_features_i).
        y: Ignored.

    Returns:
        self: Fitted estimator.

    Raises:
        ValueError: If fewer than 2 views are provided.
    """
    # Store processed tau for use in _update_weight
    self._tau: list[float] = []  # set in super().fit via _setup_fit
    super().fit(views, y)
    return self

ADMMCCA

ADMMCCA(
    latent_dimensions: int = 1,
    center: bool = True,
    tau: float | list[float] = 0.1,
    mu: float = 1.0,
    max_iter: int = 500,
    admm_iter: int = 50,
    tol: float = 1e-06,
    random_state: int | None = None,
)

Bases: _BaseIterative

Sparse CCA via linearised Alternating Direction Method of Multipliers.

Suo, Mineiro & Anandkumar (2017) pose two-view sparse CCA as, for view \(i\) fixed at the other views' current weights (i.e. \(\bar{\mathbf{s}}_{\neg i}\), the summed score of every other view),

\[ \max_{\mathbf{w}_i}\ \mathbf{w}_i^\top X_i^\top \bar{\mathbf{s}}_{\neg i} - \tau_i \|\mathbf{w}_i\|_1 \quad\text{subject to } \|X_i \mathbf{w}_i\|_2 \le 1. \]

This is a genuinely different problem from a reduced-rank-regression-style method like :class:~cca_zoo.linear.CCAR3: the objective is linear in \(\mathbf{w}_i\) (a covariance to maximise, not a residual to shrink), and the norm constraint bounds the score \(X_i\mathbf{w}_i\), not the weight vector itself -- these coincide only when \(X_i\) is orthonormal. Because the constraint couples \(\mathbf{w}_i\) to \(X_i\mathbf{w}_i\) through a linear map rather than the identity, an ordinary ADMM split would need to invert \(X_i^\top X_i\) every step; instead, introducing \(\mathbf{z}_i = X_i\mathbf{w}_i\) and linearising the augmented Lagrangian's quadratic penalty term around the current iterate turns the \(\mathbf{w}_i\)-update into a single proximal-gradient step, which is closed-form here since the linear-plus-L1 objective's proximal operator is itself just a shifted soft-threshold:

\[ \begin{aligned} \mathbf{w}_i &\leftarrow \mathcal{S}_{\eta_i \tau_i}\Bigl( \mathbf{w}_i - \eta_i \mu X_i^\top(X_i\mathbf{w}_i - \mathbf{z}_i + \boldsymbol{\xi}_i) + \eta_i X_i^\top \bar{\mathbf{s}}_{\neg i} \Bigr) \\ \mathbf{z}_i &\leftarrow \Pi_{\|\cdot\|_2 \le 1}\bigl( X_i \mathbf{w}_i + \boldsymbol{\xi}_i \bigr) \\ \boldsymbol{\xi}_i &\leftarrow \boldsymbol{\xi}_i + X_i\mathbf{w}_i - \mathbf{z}_i \end{aligned} \]

where \(\mathcal{S}_\lambda\) is the elementwise soft-threshold operator, \(\Pi_{\|\cdot\|_2 \le 1}\) projects onto the unit ball, \(\mu\) is the ADMM penalty parameter, and \(\eta_i = 1/(\mu \|X_i\|_{\mathrm{op}}^2)\) is the linearisation step size (the standard stability choice, satisfying \(\eta_i \mu \|X_i\|_{\mathrm{op}}^2 \le 1\) with equality). Verified against first-order KKT optimality conditions of the exact constrained problem (the dual variable recovered from the active constraint agrees across every nonzero coordinate to machine precision, and every zeroed coordinate's subgradient residual falls inside \([-\tau_i, \tau_i]\)).

An earlier version of this class solved a different problem entirely: a reduced-rank-regression-style loss \(\|X_i\mathbf{w}_i - \bar{\mathbf{s}}_{\neg i}\|_2^2\) with the ball constraint on \(\mathbf{w}_i\) directly, rather than the paper's linear covariance objective with the constraint on \(X_i\mathbf{w}_i\). That mismatch was not caught by comparing internal consistency alone (a previous fix corrected a scaling bug within that wrong objective and still didn't match this paper); only reading the paper's own Section 2.2 -- including its derivation of why linearisation is needed for the \(Xu=z\) constraint in particular -- surfaced it.

Multiple canonical vectors are extracted by this module's shared deflation convention (see the module docstring), not the paper's own Section 2.3 (which instead augments \(X_i\) with \(U^\top X_i^\top X_i\) rows to embed a whitened-orthogonality constraint directly into the linear map) -- the same kind of documented departure :class:SAR already makes from its own two-view paper's higher-order extension.

References

Suo, X., Mineiro, P., & Anandkumar, A. (2017). Sparse canonical correlation analysis. arXiv:1705.10865.

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent dimensions. Default is 1.

1
center bool

Whether to subtract column means. Default True.

True
tau float | list[float]

L1 regularisation weight(s). Default is 0.1.

0.1
mu float

ADMM penalty parameter. Default is 1.0.

1.0
max_iter int

Maximum outer (across-view) iterations. Default is 500.

500
admm_iter int

Maximum linearised-ADMM iterations per view, per outer iteration. Default is 50.

50
tol float

Convergence tolerance, for both the outer loop and each view's inner ADMM loop. Default is 1e-6.

1e-06
random_state int | None

Seed for reproducibility.

None

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = ADMMCCA(tau=0.1, random_state=0).fit([X1, X2])
Source code in cca_zoo/sparse/_iterative.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    tau: float | list[float] = 0.1,
    mu: float = 1.0,
    max_iter: int = 500,
    admm_iter: int = 50,
    tol: float = 1e-6,
    random_state: int | None = None,
) -> None:
    super().__init__(
        latent_dimensions=latent_dimensions,
        center=center,
        max_iter=max_iter,
        tol=tol,
        random_state=random_state,
    )
    self.tau = tau
    self.mu = mu
    self.admm_iter = admm_iter

IPLSCCA

IPLSCCA(
    latent_dimensions: int = 1,
    center: bool = True,
    alpha: float | list[float] = 0.0,
    l1_ratio: float | list[float] = 1.0,
    max_iter: int = 500,
    tol: float = 1e-06,
    random_state: int | None = None,
)

Bases: _BaseIterative

Iterative PLS with elastic net penalty on weight vectors.

Alternates between penalised regression sub-problems. For view \(i\):

\[ \hat{\mathbf{w}}_i = \arg\min_{\mathbf{w}} \frac{1}{2n} \|X_i \mathbf{w} - \bar{\mathbf{s}}_{\neg i}\|_2^2 + \alpha_i \Bigl( l_1 \|\mathbf{w}\|_1 + \tfrac{1-l_1}{2} \|\mathbf{w}\|_2^2 \Bigr) \]

followed by a normalisation step to enforce unit variance of the score.

References

Mai, Q., & Zhang, X. (2019). An iterative penalized least squares approach to sparse canonical correlation analysis. Biometrics, 75(3), 734–744.

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent dimensions. Default is 1.

1
center bool

Whether to subtract column means. Default True.

True
alpha float | list[float]

Elastic net penalty strength(s). Default is 0.

0.0
l1_ratio float | list[float]

Ratio of L1 to total penalty. 1 = lasso, 0 = ridge. Default is 1.

1.0
max_iter int

Maximum ALS iterations. Default is 500.

500
tol float

Convergence tolerance. Default is 1e-6.

1e-06
random_state int | None

Seed for reproducibility.

None

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = IPLSCCA(alpha=0.1, random_state=0).fit([X1, X2])
Source code in cca_zoo/sparse/_iterative.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    alpha: float | list[float] = 0.0,
    l1_ratio: float | list[float] = 1.0,
    max_iter: int = 500,
    tol: float = 1e-6,
    random_state: int | None = None,
) -> None:
    super().__init__(
        latent_dimensions=latent_dimensions,
        center=center,
        max_iter=max_iter,
        tol=tol,
        random_state=random_state,
    )
    self.alpha = alpha
    self.l1_ratio = l1_ratio

SpanCCA

SpanCCA(
    latent_dimensions: int = 1,
    center: bool = True,
    span: int | list[int] | None = None,
    max_iter: int = 500,
    tol: float = 1e-06,
    random_state: int | None = None,
)

Bases: _BaseIterative

Hard-thresholding ALS for sparse CCA, inspired by the SpanCCA algorithm.

Solves sparse CCA by an alternating least squares loop where each weight update retains only the span entries with the largest absolute values.

Note that this class shares its name with the paper's own algorithm (Asteris et al. 2016's "SpanCCA") but is an ALS-based heuristic, not a reimplementation of it: the paper instead takes a single rank-r SVD of the cross-covariance matrix up front, then draws many independent random directions on the low-rank subspace, hard-thresholds each one, and returns whichever independent candidate scored highest -- it never alternately refines one running weight vector the way this class (and every other class in this module, see the module docstring) does.

References

Asteris, M., Kyrillidis, A., Koyejo, O., & Poldrack, R. (2016). A simple and provable algorithm for sparse diagonal CCA. ICML, arXiv:1605.08961.

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent dimensions. Default is 1.

1
center bool

Whether to subtract column means. Default True.

True
span int | list[int] | None

Number of non-zero entries to retain per view. Either a single int or a list. Default is None (keep all — no sparsity).

None
max_iter int

Maximum ALS iterations. Default is 500.

500
tol float

Convergence tolerance. Default is 1e-6.

1e-06
random_state int | None

Seed for reproducibility.

None

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = SpanCCA(span=5, random_state=0).fit([X1, X2])
Source code in cca_zoo/sparse/_iterative.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    span: int | list[int] | None = None,
    max_iter: int = 500,
    tol: float = 1e-6,
    random_state: int | None = None,
) -> None:
    super().__init__(
        latent_dimensions=latent_dimensions,
        center=center,
        max_iter=max_iter,
        tol=tol,
        random_state=random_state,
    )
    self.span = span

WaijenborgCCA

WaijenborgCCA(
    latent_dimensions: int = 1,
    center: bool = True,
    alpha: float | list[float] = 0.0,
    l1_ratio: float | list[float] = 0.5,
    max_iter: int = 500,
    tol: float = 1e-06,
    random_state: int | None = None,
)

Bases: _BaseIterative

Elastic net regularised CCA (Waaijenborg 2008).

Named after the paper's own author rather than the generic "elastic net CCA" it used to be called, to disambiguate it from :class:~cca_zoo.sparse.ElasticNetCCA: that class optimises the actual Eckart-Young CCA loss with an elastic-net penalty (and cannot delegate to :class:~sklearn.linear_model.ElasticNet, since the EY loss's per-coordinate restriction is an exact quartic, not a quadratic -- see :func:~cca_zoo._utils._ey.coordinate_descent_ey), whereas this class alternates between plain elastic-net regression sub-problems (regressing each view's score against the sum of all other views' scores) and genuinely does delegate to sklearn's :class:~sklearn.linear_model.ElasticNet/:class:~sklearn.linear_model.Lasso/ :class:~sklearn.linear_model.Ridge via :func:_make_regressors -- a different algorithm entirely, not just a different implementation of the same one.

Alternates between elastic net regression sub-problems, regressing each view's score against the sum of all other views' scores:

\[ \hat{\mathbf{w}}_i = \arg\min_{\mathbf{w}} \frac{1}{2n} \|X_i \mathbf{w} - \mathbf{s}_{\text{all}}\|_2^2 + \alpha_i \Bigl( l_1 \|\mathbf{w}\|_1 + \tfrac{1 - l_1}{2} \|\mathbf{w}\|_2^2 \Bigr) \]

where \(\mathbf{s}_{\text{all}} = \sum_j X_j \mathbf{w}_j / \|\cdot\|\).

References

Waaijenborg, S., de Witt Hamer, P. C. V., & Zwinderman, A. H. (2008). Quantifying the association between gene expressions and DNA-markers by penalized canonical correlation analysis. Statistical Applications in Genetics and Molecular Biology, 7(1).

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent dimensions. Default is 1.

1
center bool

Whether to subtract column means. Default True.

True
alpha float | list[float]

Elastic net regularisation strength. Default is 0.

0.0
l1_ratio float | list[float]

L1 / total penalty ratio. Default is 0.5.

0.5
max_iter int

Maximum ALS iterations. Default is 500.

500
tol float

Convergence tolerance. Default is 1e-6.

1e-06
random_state int | None

Seed for reproducibility.

None

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = WaijenborgCCA(alpha=0.1, l1_ratio=0.5, random_state=0).fit([X1, X2])
Source code in cca_zoo/sparse/_iterative.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    alpha: float | list[float] = 0.0,
    l1_ratio: float | list[float] = 0.5,
    max_iter: int = 500,
    tol: float = 1e-6,
    random_state: int | None = None,
) -> None:
    super().__init__(
        latent_dimensions=latent_dimensions,
        center=center,
        max_iter=max_iter,
        tol=tol,
        random_state=random_state,
    )
    self.alpha = alpha
    self.l1_ratio = l1_ratio

ParkhomenkoCCA

ParkhomenkoCCA(
    latent_dimensions: int = 1,
    center: bool = True,
    tau: float | list[float] = 0.1,
    max_iter: int = 500,
    tol: float = 1e-06,
    random_state: int | None = None,
)

Bases: _BaseIterative

Sparse CCA via soft-thresholding power iteration (Parkhomenko 2009).

The paper's own criterion is \(K = \hat\Sigma_{XX}^{-1/2} \hat\Sigma_{XY} \hat\Sigma_{YY}^{-1/2}\), with \(\hat\Sigma_{XX}, \hat\Sigma_{YY}\) replaced by their diagonals -- the paper states directly that sparsity is only guaranteed under this diagonal approximation, not the full sample covariance. Since diagonal whitening of a matrix is exactly per-column standardisation, this is implemented by standardising each view to unit per-feature variance once per latent dimension (on top of the existing mean-centring), then running the same power iteration :class:PMDCCA's raw-covariance methods use on that standardised data, with a fixed soft-threshold \(\tau_i\) in place of the adaptive bisection search:

\[ \mathbf{w}_i \leftarrow S_{\tau_i}(\tilde X_i^\top \bar{\mathbf{s}}_{\neg i}) \]

where \(S_\tau\) is the element-wise soft-threshold operator and \(\tilde X_i\) denotes \(X_i\) with each column scaled to unit variance. The final weight is converted back to the original (unstandardised) feature scale before being returned.

References

Parkhomenko, E., Tritchler, D., & Beyene, J. (2009). Sparse canonical correlation analysis with application to genomic data integration. Statistical Applications in Genetics and Molecular Biology, 8(1).

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent dimensions. Default is 1.

1
center bool

Whether to subtract column means. Default True.

True
tau float | list[float]

Soft-threshold parameter(s). Default is 0.1.

0.1
max_iter int

Maximum ALS iterations. Default is 500.

500
tol float

Convergence tolerance. Default is 1e-6.

1e-06
random_state int | None

Seed for reproducibility.

None

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 10))
>>> X2 = rng.standard_normal((50, 8))
>>> model = ParkhomenkoCCA(tau=0.1, random_state=0).fit([X1, X2])
Source code in cca_zoo/sparse/_iterative.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    tau: float | list[float] = 0.1,
    max_iter: int = 500,
    tol: float = 1e-6,
    random_state: int | None = None,
) -> None:
    super().__init__(
        latent_dimensions=latent_dimensions,
        center=center,
        max_iter=max_iter,
        tol=tol,
        random_state=random_state,
    )
    self.tau = tau

SAR

SAR(
    latent_dimensions: int = 1,
    center: bool = True,
    n_lambda: int = 100,
    max_iter: int = 500,
    tol: float = 1e-06,
    random_state: int | None = None,
)

Bases: _BaseIterative

Sparse Alternating Regression, with BIC-selected sparsity.

Wilms \& Croux (2015) recast CCA as the Brillinger (1975) / Izenman (1975) regression problem \((\hat A, \hat B) = \arg\min_{A,B}\sum_i\|A^\top \mathbf{x}_i - B^\top\mathbf{y}_i\|_2^2\), solved by alternating regression (Wold 1968): with \(A\) fixed, \(B\) is a regression of the score \(XA\) on \(Y\); with \(B\) fixed, \(A\) is a regression of \(YB\) on \(X\). Ordinary least squares makes this exact but neither sparse nor usable once a view has more features than samples, so each regression step is replaced by a lasso fit (:func:_sar_bic_lasso) with its penalty strength selected by BIC rather than left as a user-set hyperparameter -- the paper's own point of departure from every other alternating-regression method in this module. The multiview generalisation (each view regressed against the summed score of every other view, via the same :func:_target_score helper :class:SpanCCA, :class:ParkhomenkoCCA, and :class:WaijenborgCCA use) is this implementation's own extension, not something the two-view paper itself considers.

Because a lasso fit does not commute with deflation the way an ordinary-least-squares fit does, latent dimensions beyond the first need an extra step the rest of this module's classes do not: after alternating regression on the deflated views for dimension \(d>0\) yields a direction in the deflated coordinate system, that direction's score is regressed once more against each view's original, undeflated data (again by BIC-selected lasso) to obtain the final sparse weight vector reported for that dimension -- Wilms \& Croux (2015, Section 3, "Higher order canonical vector pairs") introduce this re-expression step for exactly this reason. The first dimension needs no such step, since the deflated and original views coincide before any deflation has happened.

References

Wilms, I., & Croux, C. (2015). Sparse canonical correlation analysis from a predictive point of view. Biometrical Journal, 57(5), 834-851. arXiv:1501.01231.

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent dimensions. Default is 1.

1
center bool

Whether to subtract column means. Default True.

True
n_lambda int

Number of points in each BIC-selected lasso path's automatically-generated \(\lambda\) grid. Default is 100.

100
max_iter int

Maximum ALS iterations per latent dimension. Default is 500.

500
tol float

Convergence tolerance, for both the ALS loop and each lasso path's coordinate descent. Default is 1e-6.

1e-06
random_state int | None

Seed for reproducible random initialisation.

None

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> latent = rng.standard_normal(50)
>>> X1 = np.column_stack([latent, rng.standard_normal((50, 9))])
>>> X2 = np.column_stack([latent, rng.standard_normal((50, 7))])
>>> model = SAR(random_state=0).fit([X1, X2])
Source code in cca_zoo/sparse/_iterative.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    n_lambda: int = 100,
    max_iter: int = 500,
    tol: float = 1e-6,
    random_state: int | None = None,
) -> None:
    super().__init__(
        latent_dimensions=latent_dimensions,
        center=center,
        max_iter=max_iter,
        tol=tol,
        random_state=random_state,
    )
    self.n_lambda = n_lambda

fit

fit(views: list[ArrayLike], y: None = None) -> SAR

Fit by ALS with deflation.

Dimensions past the first are re-expressed in each view's original (undeflated) coordinates.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of arrays, each (n_samples, n_features_i).

required
y None

Ignored.

None

Returns:

Name Type Description
self SAR

Fitted estimator.

Source code in cca_zoo/sparse/_iterative.py
def fit(self, views: list[ArrayLike], y: None = None) -> SAR:
    """Fit by ALS with deflation.

    Dimensions past the first are re-expressed in each view's
    original (undeflated) coordinates.

    Args:
        views: List of arrays, each (n_samples, n_features_i).
        y: Ignored.

    Returns:
        self: Fitted estimator.
    """
    views_: list[np.ndarray] = self._setup_fit(views)
    rng = np.random.default_rng(self.random_state)
    self.weights_: list[np.ndarray] = [
        np.zeros((p, self.latent_dimensions)) for p in self.n_features_in_
    ]
    deflated = [v.copy() for v in views_]
    for d in range(self.latent_dimensions):
        w = [rng.standard_normal(p) for p in self.n_features_in_]
        w = [wi / np.linalg.norm(wi) for wi in w]
        self._fit_single(deflated, w, d)
        if d == 0:
            w_final = w
        else:
            w_final = [
                self._reexpress(views_[j], deflated[j] @ w[j])
                for j in range(self.n_views_)
            ]
        for j in range(self.n_views_):
            self.weights_[j][:, d] = w_final[j]
        deflated = deflate(deflated, w)
    return self