Skip to content

cca_zoo.stochastic

Mini-batch CCA methods, for data too large to fit a full-batch gradient step in memory.


StochasticCCAEY

StochasticCCAEY(
    latent_dimensions: int = 1,
    center: bool = True,
    c: float = 0.0,
    learning_rate: float = 0.01,
    momentum: float = 0.9,
    batch_size: int | None = None,
    max_iter: int = 1000,
    tol: float = 1e-06,
    random_state: int | None = None,
)

Bases: CCAEY

Eckart-Young CCA fit by mini-batch momentum SGD, for large-scale data.

Identical objective to :class:~cca_zoo.linear.gradient.CCAEY (same c ridge blend towards :class:~cca_zoo.linear.gradient.PLSEY, same analytic gradient) but fit the way :class:~sklearn.linear_model.SGDRegressor fits a linear model: each epoch, the data is shuffled once and split into batch_size chunks (:func:sklearn.utils.gen_batches), taking one momentum gradient step per chunk. Use this instead of :class:~cca_zoo.linear.gradient.CCAEY when the full dataset does not fit comfortably in memory or a full-batch gradient evaluation is too slow to repeat every iteration.

Parameters:

Name Type Description Default
latent_dimensions int

Number of latent dimensions. Default is 1.

1
center bool

Whether to subtract column means. Default True.

True
c float

Ridge blend in [0, 1] between CCAEY (0) and PLSEY (1). Default is 0; see :class:~cca_zoo.linear.gradient.CCAEY's docstring for the numerical-stability note on high-dimensional data.

0.0
learning_rate float

Gradient step size. Default is 1e-2.

0.01
momentum float

Momentum coefficient in [0, 1). Default is 0.9.

0.9
batch_size int | None

Mini-batch size. None uses the full dataset (one gradient step per epoch).

None
max_iter int

Number of epochs (full passes over the shuffled data). Default is 1000.

1000
tol float

Convergence tolerance on the full-dataset objective's change between consecutive epochs. Default is 1e-6.

1e-06
random_state int | None

Seed for reproducibility.

None

Examples:

>>> import numpy as np
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((5000, 200))
>>> X2 = rng.standard_normal((5000, 150))
>>> model = StochasticCCAEY(latent_dimensions=4, batch_size=128, random_state=0)
>>> model = model.fit([X1, X2])
Source code in cca_zoo/stochastic/_stochastic_cca_ey.py
def __init__(
    self,
    latent_dimensions: int = 1,
    center: bool = True,
    c: float = 0.0,
    learning_rate: float = 1e-2,
    momentum: float = 0.9,
    batch_size: int | None = None,
    max_iter: int = 1000,
    tol: float = 1e-6,
    random_state: int | None = None,
) -> None:
    super().__init__(
        latent_dimensions=latent_dimensions,
        center=center,
        c=c,
        max_iter=max_iter,
        tol=tol,
        random_state=random_state,
    )
    self.learning_rate = learning_rate
    self.momentum = momentum
    self.batch_size = batch_size

fit

fit(
    views: list[ArrayLike], y: None = None
) -> StochasticCCAEY

Fit by mini-batch momentum SGD on the EY loss.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of 2 or more arrays, each (n_samples, n_features_i).

required
y None

Ignored.

None

Returns:

Name Type Description
self StochasticCCAEY

Fitted estimator.

Raises:

Type Description
ValueError

If fewer than 2 views are provided.

ValueError

If views have inconsistent numbers of samples.

Source code in cca_zoo/stochastic/_stochastic_cca_ey.py
def fit(self, views: list[ArrayLike], y: None = None) -> StochasticCCAEY:
    """Fit by mini-batch momentum SGD on the EY loss.

    Args:
        views: List of 2 or more arrays, each (n_samples, n_features_i).
        y: Ignored.

    Returns:
        self: Fitted estimator.

    Raises:
        ValueError: If fewer than 2 views are provided.
        ValueError: If views have inconsistent numbers of samples.
    """
    views_: list[np.ndarray] = self._setup_fit(views)
    rng = np.random.default_rng(self.random_state)
    self.weights_ = self._fit_sgd(views_, rng)
    return self