Skip to content

cca_zoo.model_selection

Cross-validated hyperparameter search and significance testing for multiview models.


GridSearchCV

GridSearchCV(
    estimator: BaseEstimator,
    param_grid: dict[str, list[Any]]
    | list[dict[str, list[Any]]],
    *,
    cv: int | Any = 5,
    scoring: str | None = None,
    n_jobs: int | None = None,
    refit: bool = True,
    verbose: int = 0,
    pre_dispatch: str | int = "2*n_jobs",
    error_score: float = np.nan,
    return_train_score: bool = False,
)

Bases: _BaseMultiviewSearchCV

Exhaustive grid search with cross-validation for multiview CCA models.

A thin multiview adapter around :class:sklearn.model_selection.GridSearchCV: views are horizontally stacked into one array via :class:MultiviewWrapper, and the actual search (candidate generation, parallel fold evaluation, cv_results_, refitting, ...) is entirely sklearn's.

Parameters:

Name Type Description Default
estimator BaseEstimator

A multiview CCA estimator (e.g. :class:~cca_zoo.linear.CCA).

required
param_grid dict[str, list[Any]] | list[dict[str, list[Any]]]

Dictionary or list of dictionaries with parameter names as keys and lists of parameter settings as values.

required
cv int | Any

Number of cross-validation folds or a cross-validation splitter. Default is 5.

5
scoring str | None

Scoring strategy. When None the estimator's default :meth:score method is used.

None
n_jobs int | None

Number of jobs to run in parallel. Default is None (sequential).

None
refit bool

Whether to refit the best estimator on the full dataset. Default is True.

True
verbose int

Verbosity level. Default is 0.

0
pre_dispatch str | int

Controls the number of jobs dispatched during parallel execution, forwarded to sklearn's GridSearchCV.

'2*n_jobs'
error_score float

Value to assign to the score if fitting a candidate raises an exception, forwarded to sklearn's GridSearchCV.

nan
return_train_score bool

If True, cv_results_ also includes training scores.

False

Examples:

>>> import numpy as np
>>> from cca_zoo.linear import CCA
>>> from cca_zoo.model_selection import GridSearchCV
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 5))
>>> X2 = rng.standard_normal((50, 4))
>>> gs = GridSearchCV(
...     CCA(), param_grid={"latent_dimensions": [1, 2]}, cv=2
... )
>>> gs = gs.fit([X1, X2])

A per-view estimator parameter (e.g. :class:~cca_zoo.linear.rCCA's ridge c) can be searched independently per view with a name__<view index> suffix in param_grid:

>>> from cca_zoo.linear import rCCA
>>> gs = GridSearchCV(
...     rCCA(), param_grid={"c__0": [0.0, 0.1], "c__1": [0.0, 0.5]}, cv=2
... )
>>> gs = gs.fit([X1, X2])
>>> sorted(gs.best_params_.items())
[('c__0', 0.0), ('c__1', 0.5)]
Source code in cca_zoo/model_selection/_search.py
def __init__(
    self,
    estimator: BaseEstimator,
    param_grid: dict[str, list[Any]] | list[dict[str, list[Any]]],
    *,
    cv: int | Any = 5,
    scoring: str | None = None,
    n_jobs: int | None = None,
    refit: bool = True,
    verbose: int = 0,
    pre_dispatch: str | int = "2*n_jobs",
    error_score: float = np.nan,
    return_train_score: bool = False,
) -> None:
    self.estimator = estimator
    self.param_grid = param_grid
    self.cv = cv
    self.scoring = scoring
    self.n_jobs = n_jobs
    self.refit = refit
    self.verbose = verbose
    self.pre_dispatch = pre_dispatch
    self.error_score = error_score
    self.return_train_score = return_train_score

fit

fit(
    views: list[ArrayLike],
    y: None = None,
    **fit_params: Any,
) -> GridSearchCV

Run grid search with cross-validation on multiview data.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of arrays, each of shape (n_samples, n_features_i). All arrays must have the same number of rows.

required
y None

Ignored.

None
**fit_params Any

Additional keyword arguments forwarded to the estimator's fit method during each fold.

{}

Returns:

Name Type Description
self GridSearchCV

Fitted grid search object.

Source code in cca_zoo/model_selection/_search.py
def fit(
    self,
    views: list[ArrayLike],
    y: None = None,
    **fit_params: Any,
) -> GridSearchCV:
    """Run grid search with cross-validation on multiview data.

    Args:
        views: List of arrays, each of shape (n_samples, n_features_i).
            All arrays must have the same number of rows.
        y: Ignored.
        **fit_params: Additional keyword arguments forwarded to the
            estimator's ``fit`` method during each fold.

    Returns:
        self: Fitted grid search object.
    """
    inner_cv_kwargs = dict(
        param_grid=_wrap_param_space(self.param_grid),
        cv=self.cv,
        scoring=self.scoring,
        n_jobs=self.n_jobs,
        refit=self.refit,
        verbose=self.verbose,
        pre_dispatch=self.pre_dispatch,
        error_score=self.error_score,
        return_train_score=self.return_train_score,
    )
    return cast("GridSearchCV", self._fit(views, y, inner_cv_kwargs, **fit_params))

RandomizedSearchCV

RandomizedSearchCV(
    estimator: BaseEstimator,
    param_distributions: dict[str, Any]
    | list[dict[str, Any]],
    *,
    n_iter: int = 10,
    cv: int | Any = 5,
    scoring: str | None = None,
    n_jobs: int | None = None,
    refit: bool = True,
    verbose: int = 0,
    random_state: int | Any = None,
    pre_dispatch: str | int = "2*n_jobs",
    error_score: float = np.nan,
    return_train_score: bool = False,
)

Bases: _BaseMultiviewSearchCV

Randomized search with cross-validation for multiview CCA models.

Samples n_iter parameter settings from param_distributions instead of exhaustively trying every combination in a grid -- useful when a hyperparameter (e.g. c) is continuous, or when the grid is too large to search exhaustively. A thin multiview adapter around :class:sklearn.model_selection.RandomizedSearchCV, following the same :class:MultiviewWrapper pattern as :class:GridSearchCV.

Parameters:

Name Type Description Default
estimator BaseEstimator

A multiview CCA estimator (e.g. :class:~cca_zoo.linear.CCA).

required
param_distributions dict[str, Any] | list[dict[str, Any]]

Dictionary (or list of dictionaries) with parameter names as keys and either a list of values to sample from, or a distribution (anything with a rvs method, e.g. scipy.stats.loguniform).

required
n_iter int

Number of parameter settings sampled. Default is 10.

10
cv int | Any

Number of cross-validation folds or a cross-validation splitter. Default is 5.

5
scoring str | None

Scoring strategy. When None the estimator's default :meth:score method is used.

None
n_jobs int | None

Number of jobs to run in parallel. Default is None (sequential).

None
refit bool

Whether to refit the best estimator on the full dataset. Default is True.

True
verbose int

Verbosity level. Default is 0.

0
random_state int | Any

Controls the randomness of the parameter sampling.

None
pre_dispatch str | int

Controls the number of jobs dispatched during parallel execution, forwarded to sklearn's RandomizedSearchCV.

'2*n_jobs'
error_score float

Value to assign to the score if fitting a candidate raises an exception, forwarded to sklearn's RandomizedSearchCV.

nan
return_train_score bool

If True, cv_results_ also includes training scores.

False

Examples:

>>> import numpy as np
>>> from scipy.stats import loguniform
>>> from cca_zoo.linear import rCCA
>>> from cca_zoo.model_selection import RandomizedSearchCV
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 5))
>>> X2 = rng.standard_normal((50, 4))
>>> rs = RandomizedSearchCV(
...     rCCA(),
...     param_distributions={"c": loguniform(1e-3, 1.0)},
...     n_iter=5,
...     cv=2,
...     random_state=0,
... )
>>> rs = rs.fit([X1, X2])

Per-view distributions use the same name__<view index> suffix as :class:GridSearchCV:

>>> rs = RandomizedSearchCV(
...     rCCA(),
...     param_distributions={
...         "c__0": loguniform(1e-3, 1.0),
...         "c__1": loguniform(1e-3, 1.0),
...     },
...     n_iter=5,
...     cv=2,
...     random_state=0,
... )
>>> rs = rs.fit([X1, X2])
>>> sorted(rs.best_params_.keys())
['c__0', 'c__1']
Source code in cca_zoo/model_selection/_search.py
def __init__(
    self,
    estimator: BaseEstimator,
    param_distributions: dict[str, Any] | list[dict[str, Any]],
    *,
    n_iter: int = 10,
    cv: int | Any = 5,
    scoring: str | None = None,
    n_jobs: int | None = None,
    refit: bool = True,
    verbose: int = 0,
    random_state: int | Any = None,
    pre_dispatch: str | int = "2*n_jobs",
    error_score: float = np.nan,
    return_train_score: bool = False,
) -> None:
    self.estimator = estimator
    self.param_distributions = param_distributions
    self.n_iter = n_iter
    self.cv = cv
    self.scoring = scoring
    self.n_jobs = n_jobs
    self.refit = refit
    self.verbose = verbose
    self.random_state = random_state
    self.pre_dispatch = pre_dispatch
    self.error_score = error_score
    self.return_train_score = return_train_score

fit

fit(
    views: list[ArrayLike],
    y: None = None,
    **fit_params: Any,
) -> RandomizedSearchCV

Run randomized search with cross-validation on multiview data.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of arrays, each of shape (n_samples, n_features_i). All arrays must have the same number of rows.

required
y None

Ignored.

None
**fit_params Any

Additional keyword arguments forwarded to the estimator's fit method during each fold.

{}

Returns:

Name Type Description
self RandomizedSearchCV

Fitted randomized search object.

Source code in cca_zoo/model_selection/_search.py
def fit(
    self,
    views: list[ArrayLike],
    y: None = None,
    **fit_params: Any,
) -> RandomizedSearchCV:
    """Run randomized search with cross-validation on multiview data.

    Args:
        views: List of arrays, each of shape (n_samples, n_features_i).
            All arrays must have the same number of rows.
        y: Ignored.
        **fit_params: Additional keyword arguments forwarded to the
            estimator's ``fit`` method during each fold.

    Returns:
        self: Fitted randomized search object.
    """
    inner_cv_kwargs = dict(
        param_distributions=_wrap_param_space(self.param_distributions),
        n_iter=self.n_iter,
        cv=self.cv,
        scoring=self.scoring,
        n_jobs=self.n_jobs,
        refit=self.refit,
        verbose=self.verbose,
        random_state=self.random_state,
        pre_dispatch=self.pre_dispatch,
        error_score=self.error_score,
        return_train_score=self.return_train_score,
    )
    return cast(
        "RandomizedSearchCV", self._fit(views, y, inner_cv_kwargs, **fit_params)
    )

HalvingGridSearchCV

HalvingGridSearchCV(
    estimator: BaseEstimator,
    param_grid: dict[str, list[Any]]
    | list[dict[str, list[Any]]],
    *,
    factor: int | float = 3,
    resource: str = "n_samples",
    max_resources: int | str = "auto",
    min_resources: int | str = "exhaust",
    aggressive_elimination: bool = False,
    cv: int | Any = 5,
    scoring: str | None = None,
    refit: bool = True,
    error_score: float = np.nan,
    return_train_score: bool = True,
    random_state: int | Any = None,
    n_jobs: int | None = None,
    verbose: int = 0,
)

Bases: _BaseMultiviewSearchCV

Successive-halving grid search with cross-validation for multiview CCA models.

Like :class:GridSearchCV, but candidates are evaluated on a growing subset of the training samples across rounds: most candidates are eliminated early on a small subset, and only the survivors are evaluated on progressively larger subsets, which is usually much cheaper than an exhaustive :class:GridSearchCV when the grid is large. A thin multiview adapter around :class:sklearn.model_selection.HalvingGridSearchCV, following the same :class:MultiviewWrapper pattern as :class:GridSearchCV; the "resource" being grown across rounds is a row count of the (already view-concatenated) training array, so the successive-halving mechanics need no multiview-specific handling.

Parameters:

Name Type Description Default
estimator BaseEstimator

A multiview CCA estimator (e.g. :class:~cca_zoo.linear.CCA).

required
param_grid dict[str, list[Any]] | list[dict[str, list[Any]]]

Dictionary or list of dictionaries with parameter names as keys and lists of parameter settings as values.

required
factor int | float

The proportion of candidates eliminated (and resources multiplied by) at each round. Default is 3.

3
resource str

The resource grown between rounds, forwarded to sklearn's HalvingGridSearchCV. Default is "n_samples".

'n_samples'
max_resources int | str

The maximum amount of resource a candidate is allowed to use, forwarded to sklearn's HalvingGridSearchCV. Default is "auto".

'auto'
min_resources int | str

The minimum amount of resource a candidate is allowed to use, forwarded to sklearn's HalvingGridSearchCV. Default is "exhaust".

'exhaust'
aggressive_elimination bool

Whether to eliminate candidates at the same rate even before there are enough resources to grow, forwarded to sklearn's HalvingGridSearchCV. Default is False.

False
cv int | Any

Number of cross-validation folds or a cross-validation splitter. Default is 5.

5
scoring str | None

Scoring strategy. When None the estimator's default :meth:score method is used.

None
refit bool

Whether to refit the best estimator on the full dataset. Default is True.

True
error_score float

Value to assign to the score if fitting a candidate raises an exception, forwarded to sklearn's HalvingGridSearchCV.

nan
return_train_score bool

If True, cv_results_ also includes training scores. Default is True, matching sklearn's HalvingGridSearchCV.

True
random_state int | Any

Controls the pseudo-random subsampling of the training set that determines the candidates' resources at each round.

None
n_jobs int | None

Number of jobs to run in parallel. Default is None (sequential).

None
verbose int

Verbosity level. Default is 0.

0

Examples:

>>> import numpy as np
>>> from cca_zoo.linear import CCA
>>> from cca_zoo.model_selection import HalvingGridSearchCV
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 5))
>>> X2 = rng.standard_normal((50, 4))
>>> hgs = HalvingGridSearchCV(
...     CCA(), param_grid={"latent_dimensions": [1, 2]}, cv=2
... )
>>> hgs = hgs.fit([X1, X2])

Per-view parameters use the same name__<view index> suffix as :class:GridSearchCV:

>>> from cca_zoo.linear import rCCA
>>> hgs = HalvingGridSearchCV(
...     rCCA(),
...     param_grid={"c__0": [0.0, 0.1], "c__1": [0.0, 0.5]},
...     cv=2,
...     random_state=0,
... )
>>> hgs = hgs.fit([X1, X2])
>>> sorted(hgs.best_params_.items())
[('c__0', 0.1), ('c__1', 0.0)]
Source code in cca_zoo/model_selection/_search.py
def __init__(
    self,
    estimator: BaseEstimator,
    param_grid: dict[str, list[Any]] | list[dict[str, list[Any]]],
    *,
    factor: int | float = 3,
    resource: str = "n_samples",
    max_resources: int | str = "auto",
    min_resources: int | str = "exhaust",
    aggressive_elimination: bool = False,
    cv: int | Any = 5,
    scoring: str | None = None,
    refit: bool = True,
    error_score: float = np.nan,
    return_train_score: bool = True,
    random_state: int | Any = None,
    n_jobs: int | None = None,
    verbose: int = 0,
) -> None:
    self.estimator = estimator
    self.param_grid = param_grid
    self.factor = factor
    self.resource = resource
    self.max_resources = max_resources
    self.min_resources = min_resources
    self.aggressive_elimination = aggressive_elimination
    self.cv = cv
    self.scoring = scoring
    self.refit = refit
    self.error_score = error_score
    self.return_train_score = return_train_score
    self.random_state = random_state
    self.n_jobs = n_jobs
    self.verbose = verbose

fit

fit(
    views: list[ArrayLike],
    y: None = None,
    **fit_params: Any,
) -> HalvingGridSearchCV

Run successive-halving grid search with cross-validation.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of arrays, each of shape (n_samples, n_features_i). All arrays must have the same number of rows.

required
y None

Ignored.

None
**fit_params Any

Additional keyword arguments forwarded to the estimator's fit method during each fold.

{}

Returns:

Name Type Description
self HalvingGridSearchCV

Fitted search object.

Source code in cca_zoo/model_selection/_search.py
def fit(
    self,
    views: list[ArrayLike],
    y: None = None,
    **fit_params: Any,
) -> HalvingGridSearchCV:
    """Run successive-halving grid search with cross-validation.

    Args:
        views: List of arrays, each of shape (n_samples, n_features_i).
            All arrays must have the same number of rows.
        y: Ignored.
        **fit_params: Additional keyword arguments forwarded to the
            estimator's ``fit`` method during each fold.

    Returns:
        self: Fitted search object.
    """
    inner_cv_kwargs = dict(
        param_grid=_wrap_param_space(self.param_grid),
        factor=self.factor,
        resource=self.resource,
        max_resources=self.max_resources,
        min_resources=self.min_resources,
        aggressive_elimination=self.aggressive_elimination,
        cv=self.cv,
        scoring=self.scoring,
        refit=self.refit,
        error_score=self.error_score,
        return_train_score=self.return_train_score,
        random_state=self.random_state,
        n_jobs=self.n_jobs,
        verbose=self.verbose,
    )
    return cast(
        "HalvingGridSearchCV", self._fit(views, y, inner_cv_kwargs, **fit_params)
    )

HalvingRandomSearchCV

HalvingRandomSearchCV(
    estimator: BaseEstimator,
    param_distributions: dict[str, Any]
    | list[dict[str, Any]],
    *,
    n_candidates: int | str = "exhaust",
    factor: int | float = 3,
    resource: str = "n_samples",
    max_resources: int | str = "auto",
    min_resources: int | str = "smallest",
    aggressive_elimination: bool = False,
    cv: int | Any = 5,
    scoring: str | None = None,
    refit: bool = True,
    error_score: float = np.nan,
    return_train_score: bool = True,
    random_state: int | Any = None,
    n_jobs: int | None = None,
    verbose: int = 0,
)

Bases: _BaseMultiviewSearchCV

Successive-halving randomized search with cross-validation for multiview models.

Combines :class:RandomizedSearchCV's sampling of param_distributions with :class:HalvingGridSearchCV's successive-halving elimination: most sampled candidates are eliminated early on a small subset of the training samples, and only the survivors are evaluated on progressively larger subsets. A thin multiview adapter around :class:sklearn.model_selection.HalvingRandomSearchCV, following the same :class:MultiviewWrapper pattern as :class:RandomizedSearchCV.

Parameters:

Name Type Description Default
estimator BaseEstimator

A multiview CCA estimator (e.g. :class:~cca_zoo.linear.CCA).

required
param_distributions dict[str, Any] | list[dict[str, Any]]

Dictionary (or list of dictionaries) with parameter names as keys and either a list of values to sample from, or a distribution (anything with a rvs method, e.g. scipy.stats.loguniform).

required
n_candidates int | str

The number of candidate parameters to sample, forwarded to sklearn's HalvingRandomSearchCV. Default is "exhaust".

'exhaust'
factor int | float

The proportion of candidates eliminated (and resources multiplied by) at each round. Default is 3.

3
resource str

The resource grown between rounds, forwarded to sklearn's HalvingRandomSearchCV. Default is "n_samples".

'n_samples'
max_resources int | str

The maximum amount of resource a candidate is allowed to use, forwarded to sklearn's HalvingRandomSearchCV. Default is "auto".

'auto'
min_resources int | str

The minimum amount of resource a candidate is allowed to use, forwarded to sklearn's HalvingRandomSearchCV. Default is "smallest".

'smallest'
aggressive_elimination bool

Whether to eliminate candidates at the same rate even before there are enough resources to grow, forwarded to sklearn's HalvingRandomSearchCV. Default is False.

False
cv int | Any

Number of cross-validation folds or a cross-validation splitter. Default is 5.

5
scoring str | None

Scoring strategy. When None the estimator's default :meth:score method is used.

None
refit bool

Whether to refit the best estimator on the full dataset. Default is True.

True
error_score float

Value to assign to the score if fitting a candidate raises an exception, forwarded to sklearn's HalvingRandomSearchCV.

nan
return_train_score bool

If True, cv_results_ also includes training scores. Default is True, matching sklearn's HalvingRandomSearchCV.

True
random_state int | Any

Controls both the randomness of the parameter sampling and the pseudo-random subsampling of the training set.

None
n_jobs int | None

Number of jobs to run in parallel. Default is None (sequential).

None
verbose int

Verbosity level. Default is 0.

0

Examples:

>>> import numpy as np
>>> from scipy.stats import loguniform
>>> from cca_zoo.linear import rCCA
>>> from cca_zoo.model_selection import HalvingRandomSearchCV
>>> rng = np.random.default_rng(0)
>>> X1 = rng.standard_normal((50, 5))
>>> X2 = rng.standard_normal((50, 4))
>>> hrs = HalvingRandomSearchCV(
...     rCCA(),
...     param_distributions={"c": loguniform(1e-3, 1.0)},
...     cv=2,
...     random_state=0,
... )
>>> hrs = hrs.fit([X1, X2])

Per-view distributions use the same name__<view index> suffix as :class:RandomizedSearchCV:

>>> hrs = HalvingRandomSearchCV(
...     rCCA(),
...     param_distributions={
...         "c__0": loguniform(1e-3, 1.0),
...         "c__1": loguniform(1e-3, 1.0),
...     },
...     cv=2,
...     random_state=0,
... )
>>> hrs = hrs.fit([X1, X2])
>>> sorted(hrs.best_params_.keys())
['c__0', 'c__1']
Source code in cca_zoo/model_selection/_search.py
def __init__(
    self,
    estimator: BaseEstimator,
    param_distributions: dict[str, Any] | list[dict[str, Any]],
    *,
    n_candidates: int | str = "exhaust",
    factor: int | float = 3,
    resource: str = "n_samples",
    max_resources: int | str = "auto",
    min_resources: int | str = "smallest",
    aggressive_elimination: bool = False,
    cv: int | Any = 5,
    scoring: str | None = None,
    refit: bool = True,
    error_score: float = np.nan,
    return_train_score: bool = True,
    random_state: int | Any = None,
    n_jobs: int | None = None,
    verbose: int = 0,
) -> None:
    self.estimator = estimator
    self.param_distributions = param_distributions
    self.n_candidates = n_candidates
    self.factor = factor
    self.resource = resource
    self.max_resources = max_resources
    self.min_resources = min_resources
    self.aggressive_elimination = aggressive_elimination
    self.cv = cv
    self.scoring = scoring
    self.refit = refit
    self.error_score = error_score
    self.return_train_score = return_train_score
    self.random_state = random_state
    self.n_jobs = n_jobs
    self.verbose = verbose

fit

fit(
    views: list[ArrayLike],
    y: None = None,
    **fit_params: Any,
) -> HalvingRandomSearchCV

Run successive-halving randomized search with cross-validation.

Parameters:

Name Type Description Default
views list[ArrayLike]

List of arrays, each of shape (n_samples, n_features_i). All arrays must have the same number of rows.

required
y None

Ignored.

None
**fit_params Any

Additional keyword arguments forwarded to the estimator's fit method during each fold.

{}

Returns:

Name Type Description
self HalvingRandomSearchCV

Fitted search object.

Source code in cca_zoo/model_selection/_search.py
def fit(
    self,
    views: list[ArrayLike],
    y: None = None,
    **fit_params: Any,
) -> HalvingRandomSearchCV:
    """Run successive-halving randomized search with cross-validation.

    Args:
        views: List of arrays, each of shape (n_samples, n_features_i).
            All arrays must have the same number of rows.
        y: Ignored.
        **fit_params: Additional keyword arguments forwarded to the
            estimator's ``fit`` method during each fold.

    Returns:
        self: Fitted search object.
    """
    inner_cv_kwargs = dict(
        param_distributions=_wrap_param_space(self.param_distributions),
        n_candidates=self.n_candidates,
        factor=self.factor,
        resource=self.resource,
        max_resources=self.max_resources,
        min_resources=self.min_resources,
        aggressive_elimination=self.aggressive_elimination,
        cv=self.cv,
        scoring=self.scoring,
        refit=self.refit,
        error_score=self.error_score,
        return_train_score=self.return_train_score,
        random_state=self.random_state,
        n_jobs=self.n_jobs,
        verbose=self.verbose,
    )
    return cast(
        "HalvingRandomSearchCV", self._fit(views, y, inner_cv_kwargs, **fit_params)
    )

MultiviewWrapper

MultiviewWrapper(
    estimator: BaseEstimator, split_indices: list[int]
)

Bases: BaseEstimator

Adapt a multiview estimator to sklearn's single-X estimator API.

Sklearn's model-selection tools (GridSearchCV, cross_val_score, Pipeline, ...) require an estimator whose fit/score accept (X, y) with X a single 2-D array, so they can index and split rows into folds. This wrapper concatenates the views' feature axes into one array on the way in and splits them back into views on the way out, so any sklearn tool that only ever sees the concatenated array can be used unmodified with a cca_zoo multiview estimator.

Parameters:

Name Type Description Default
estimator BaseEstimator

A multiview CCA estimator (e.g. :class:~cca_zoo.linear.CCA).

required
split_indices list[int]

Number of features in each view, in order. Used to split the concatenated array back into views.

required

Examples:

>>> import numpy as np
>>> from sklearn.model_selection import cross_val_score
>>> from cca_zoo.linear import CCA
>>> from cca_zoo.model_selection import MultiviewWrapper
>>> rng = np.random.default_rng(0)
>>> X1, X2 = rng.standard_normal((50, 5)), rng.standard_normal((50, 4))
>>> wrapper = MultiviewWrapper(CCA(), split_indices=[5, 4])
>>> scores = cross_val_score(wrapper, np.hstack([X1, X2]), cv=3)

Per-view hyperparameters (a per-view CCA model accepts a scalar, broadcast to every view, or an explicit list with one value per view -- e.g. KCCA(c=[0.01, 0.1])) can be searched independently per view through set_params using a name__<view index> suffix, e.g. estimator__c__0. This is mainly useful for grid/randomized search: a param grid of {"c__0": [0.01, 0.1], "c__1": [0.1, 1.0]} makes sklearn's ParameterGrid search the Cartesian product of the two views' values. Indices left unset keep the estimator's current value for that view (broadcast if it was a scalar).

Source code in cca_zoo/model_selection/_search.py
def __init__(self, estimator: BaseEstimator, split_indices: list[int]) -> None:
    self.estimator = estimator
    self.split_indices = split_indices

set_params

set_params(**params: Any) -> MultiviewWrapper

Set parameters, honouring per-view name__<view index> overrides.

Parameters:

Name Type Description Default
**params Any

Parameter names/values. A key of the form estimator__<name>__<index> sets view <index>'s value of the wrapped estimator's <name> parameter without disturbing the other views.

{}
Source code in cca_zoo/model_selection/_search.py
def set_params(self, **params: Any) -> MultiviewWrapper:
    """Set parameters, honouring per-view ``name__<view index>`` overrides.

    Args:
        **params: Parameter names/values. A key of the form
            ``estimator__<name>__<index>`` sets view ``<index>``'s
            value of the wrapped estimator's ``<name>`` parameter
            without disturbing the other views.
    """
    own: dict[str, Any] = {}
    inner: dict[str, Any] = {}
    for key, value in params.items():
        if key.startswith(_PARAM_PREFIX):
            inner[key[len(_PARAM_PREFIX) :]] = value
        else:
            own[key] = value
    if own:
        super().set_params(**own)
    if inner:
        self._set_inner_params(**inner)
    return self

fit

fit(
    X: ndarray, y: None = None, **fit_params: Any
) -> MultiviewWrapper

Fit the wrapped estimator on the concatenated multiview data.

Source code in cca_zoo/model_selection/_search.py
def fit(self, X: np.ndarray, y: None = None, **fit_params: Any) -> MultiviewWrapper:
    """Fit the wrapped estimator on the concatenated multiview data."""
    self.estimator_ = clone(self.estimator)
    self.estimator_.fit(self._split_views(X), **fit_params)
    return self

score

score(X: ndarray, y: None = None) -> float

Mean canonical correlation over all latent dimensions.

Source code in cca_zoo/model_selection/_search.py
def score(self, X: np.ndarray, y: None = None) -> float:
    """Mean canonical correlation over all latent dimensions."""
    scores: np.ndarray = self.estimator_.score(self._split_views(X))
    return float(scores.mean())

transform

transform(X: ndarray) -> np.ndarray

Transform and re-concatenate, so the wrapper composes with Pipeline.

Source code in cca_zoo/model_selection/_search.py
def transform(self, X: np.ndarray) -> np.ndarray:
    """Transform and re-concatenate, so the wrapper composes with Pipeline."""
    transformed = self.estimator_.transform(self._split_views(X))
    return np.hstack(transformed)

permutation_test_significance

permutation_test_significance(
    estimator: BaseEstimator,
    views: list[ArrayLike],
    n_permutations: int = 1000,
    random_state: int | Generator | None = None,
    n_jobs: int | None = None,
) -> PermutationTestResult

Permutation test for canonical correlation and feature-loading significance.

Fits a clone of estimator on views, then repeatedly refits a fresh clone on data where every view except the first has had its rows independently shuffled -- destroying the true cross-view correspondence while preserving each view's own covariance structure -- to build a null distribution.

Canonical-correlation significance (p_values_) compares each dimension's observed correlation directly to its permuted counterparts: since both the observed and permuted fits rank dimensions by correlation strength, the d-th dimension of a permuted fit is already the right null comparison for the d-th observed dimension, with no realignment needed.

Feature-loading significance (loading_p_values_) is subtler: a permuted refit is not guaranteed to recover canonical variates in the same order or with the same sign as the observed fit, since permutation can induce an arbitrary rotation or reflection of near-tied dimensions (Xia et al., 2018, Nat. Commun.; McIntosh & Lobaugh, 2004, NeuroImage). Each permutation's loadings are therefore realigned to the observed loadings via :func:procrustes_rotation (fit jointly across all views' stacked loadings, since the rotation ambiguity is shared across views) before being compared feature-by-feature and dimension-by-dimension.

Parameters:

Name Type Description Default
estimator BaseEstimator

An unfitted multiview CCA/PLS estimator implementing the :class:~cca_zoo._base.BaseModel interface.

required
views list[ArrayLike]

List of arrays, each of shape (n_samples, n_features_i).

required
n_permutations int

Number of permutations to draw. Default 1000.

1000
random_state int | Generator | None

Seed or numpy.random.Generator for reproducible permutations.

None
n_jobs int | None

Number of permutations to fit in parallel (forwarded to :class:joblib.Parallel via sklearn.utils.parallel). Default None (sequential).

None

Returns:

Type Description
PermutationTestResult

PermutationTestResult with the observed and null statistics.

Raises:

Type Description
ValueError

If fewer than 2 views are provided, or n_permutations is not positive.

Examples:

>>> import numpy as np
>>> from cca_zoo.linear import CCA
>>> from cca_zoo.model_selection import permutation_test_significance
>>> rng = np.random.default_rng(0)
>>> z = rng.standard_normal((40, 1))
>>> X1 = z @ rng.standard_normal((1, 5)) + 0.1 * rng.standard_normal((40, 5))
>>> X2 = z @ rng.standard_normal((1, 4)) + 0.1 * rng.standard_normal((40, 4))
>>> result = permutation_test_significance(
...     CCA(latent_dimensions=1), [X1, X2], n_permutations=49, random_state=0
... )
>>> result.p_values_.shape
(1,)
Source code in cca_zoo/model_selection/_significance.py
def permutation_test_significance(
    estimator: BaseEstimator,
    views: list[ArrayLike],
    n_permutations: int = 1000,
    random_state: int | np.random.Generator | None = None,
    n_jobs: int | None = None,
) -> PermutationTestResult:
    """Permutation test for canonical correlation and feature-loading significance.

    Fits a clone of ``estimator`` on ``views``, then repeatedly refits a
    fresh clone on data where every view except the first has had its rows
    *independently* shuffled -- destroying the true cross-view
    correspondence while preserving each view's own covariance structure --
    to build a null distribution.

    Canonical-correlation significance (``p_values_``) compares each
    dimension's observed correlation directly to its permuted
    counterparts: since both the observed and permuted fits rank
    dimensions by correlation strength, the d-th dimension of a permuted
    fit is already the right null comparison for the d-th observed
    dimension, with no realignment needed.

    Feature-loading significance (``loading_p_values_``) is subtler: a
    permuted refit is not guaranteed to recover canonical variates in the
    same order or with the same sign as the observed fit, since
    permutation can induce an arbitrary rotation or reflection of
    near-tied dimensions (Xia et al., 2018, *Nat. Commun.*; McIntosh &
    Lobaugh, 2004, *NeuroImage*). Each permutation's loadings are
    therefore realigned to the observed loadings via
    :func:`procrustes_rotation` (fit jointly across all views' stacked
    loadings, since the rotation ambiguity is shared across views) before
    being compared feature-by-feature and dimension-by-dimension.

    Args:
        estimator: An unfitted multiview CCA/PLS estimator implementing
            the :class:`~cca_zoo._base.BaseModel` interface.
        views: List of arrays, each of shape (n_samples, n_features_i).
        n_permutations: Number of permutations to draw. Default 1000.
        random_state: Seed or ``numpy.random.Generator`` for reproducible
            permutations.
        n_jobs: Number of permutations to fit in parallel (forwarded to
            :class:`joblib.Parallel` via ``sklearn.utils.parallel``).
            Default ``None`` (sequential).

    Returns:
        PermutationTestResult with the observed and null statistics.

    Raises:
        ValueError: If fewer than 2 views are provided, or
            ``n_permutations`` is not positive.

    Examples:
        >>> import numpy as np
        >>> from cca_zoo.linear import CCA
        >>> from cca_zoo.model_selection import permutation_test_significance
        >>> rng = np.random.default_rng(0)
        >>> z = rng.standard_normal((40, 1))
        >>> X1 = z @ rng.standard_normal((1, 5)) + 0.1 * rng.standard_normal((40, 5))
        >>> X2 = z @ rng.standard_normal((1, 4)) + 0.1 * rng.standard_normal((40, 4))
        >>> result = permutation_test_significance(
        ...     CCA(latent_dimensions=1), [X1, X2], n_permutations=49, random_state=0
        ... )
        >>> result.p_values_.shape
        (1,)
    """
    arrays = validate_views(views)
    if n_permutations < 1:
        raise ValueError(f"n_permutations must be positive, got {n_permutations}.")
    n_views = len(arrays)

    fitted = clone(estimator).fit(arrays)
    true_corr = np.asarray(fitted.score(arrays))
    true_loadings = fitted.get_factor_loadings(arrays)
    true_stack = np.vstack(true_loadings)  # (sum(n_features_i), k)
    split_points = np.cumsum([loading.shape[0] for loading in true_loadings[:-1]])

    rng = np.random.default_rng(random_state)
    seeds = rng.integers(0, np.iinfo(np.int32).max, size=n_permutations)

    def _one_permutation(seed: int) -> tuple[np.ndarray, np.ndarray]:
        local_rng = np.random.default_rng(seed)
        permuted = [arrays[0]] + [local_rng.permutation(v, axis=0) for v in arrays[1:]]
        model = clone(estimator).fit(permuted)
        corr = np.asarray(model.score(permuted))
        loadings = model.get_factor_loadings(permuted)
        stack = np.vstack(loadings)
        rotation = procrustes_rotation(true_stack, stack)
        aligned_stack: np.ndarray = stack @ rotation
        return corr, aligned_stack

    results = Parallel(n_jobs=n_jobs)(delayed(_one_permutation)(seed) for seed in seeds)
    null_correlations = np.stack([corr for corr, _ in results])  # (n_perm, k)
    null_stacks = np.stack([stack for _, stack in results])  # (n_perm, sum_p, k)
    null_loadings = list(np.split(null_stacks, split_points, axis=1))

    p_values = (1 + (null_correlations >= true_corr).sum(axis=0)) / (1 + n_permutations)
    loading_p_values = [
        (1 + (np.abs(null_loadings[i]) >= np.abs(true_loadings[i])).sum(axis=0))
        / (1 + n_permutations)
        for i in range(n_views)
    ]

    return PermutationTestResult(
        correlations_=true_corr,
        null_correlations_=null_correlations,
        p_values_=p_values,
        loadings_=true_loadings,
        null_loadings_=null_loadings,
        loading_p_values_=loading_p_values,
    )

PermutationTestResult dataclass

PermutationTestResult(
    correlations_: ndarray,
    null_correlations_: ndarray,
    p_values_: ndarray,
    loadings_: list[ndarray],
    null_loadings_: list[ndarray],
    loading_p_values_: list[ndarray],
)

Result of :func:permutation_test_significance.

Attributes:

Name Type Description
correlations_ ndarray

Observed per-dimension average pairwise canonical correlations, shape (k,).

null_correlations_ ndarray

Per-dimension correlations from each permutation, shape (n_permutations, k).

p_values_ ndarray

Per-dimension permutation p-value for the canonical correlations, shape (k,).

loadings_ list[ndarray]

Observed factor loadings, one array of shape (n_features_i, k) per view (see :meth:~cca_zoo._base.BaseModel.get_factor_loadings).

null_loadings_ list[ndarray]

Permuted factor loadings, realigned to loadings_ via :func:procrustes_rotation, one array of shape (n_permutations, n_features_i, k) per view.

loading_p_values_ list[ndarray]

Per-feature, per-dimension permutation p-value for the factor loadings, one array of shape (n_features_i, k) per view.

procrustes_rotation

procrustes_rotation(
    reference: ndarray, target: ndarray
) -> np.ndarray

Solve the orthogonal Procrustes problem aligning target to reference.

Finds the orthogonal matrix R minimising ||reference - target @ R|| (Frobenius norm), via the classical SVD solution (Schönemann, 1966): writing the SVD of target.T @ reference as U @ S @ Vt, the optimum is R = U @ Vt. R is a general orthogonal matrix, not restricted to a proper (determinant +1) rotation, so it also captures axis reflections (sign flips) -- both are needed when matching permutation- or bootstrap-resampled canonical variates back to a reference fit, since resampling can induce either (Xia et al., 2018, Nat. Commun.; McIntosh & Lobaugh, 2004, NeuroImage).

Parameters:

Name Type Description Default
reference ndarray

Array of shape (n, k).

required
target ndarray

Array of shape (n, k), matched row-for-row with reference (e.g. the same features/variables in the same order), but not necessarily in the same column (component) order or sign.

required

Returns:

Type Description
ndarray

Orthogonal matrix of shape (k, k) such that target @ R is

ndarray

optimally aligned to reference in the least-squares sense.

Raises:

Type Description
ValueError

If reference and target don't have the same shape.

Examples:

>>> import numpy as np
>>> from cca_zoo.model_selection import procrustes_rotation
>>> rng = np.random.default_rng(0)
>>> reference = rng.standard_normal((20, 3))
>>> true_rotation, _ = np.linalg.qr(rng.standard_normal((3, 3)))
>>> target = reference @ true_rotation.T
>>> recovered = procrustes_rotation(reference, target)
>>> np.allclose(target @ recovered, reference, atol=1e-8)
True
Source code in cca_zoo/model_selection/_significance.py
def procrustes_rotation(reference: np.ndarray, target: np.ndarray) -> np.ndarray:
    """Solve the orthogonal Procrustes problem aligning ``target`` to ``reference``.

    Finds the orthogonal matrix ``R`` minimising ``||reference - target @
    R||`` (Frobenius norm), via the classical SVD solution (Schönemann,
    1966): writing the SVD of ``target.T @ reference`` as ``U @ S @ Vt``,
    the optimum is ``R = U @ Vt``. ``R`` is a general orthogonal matrix, not
    restricted to a proper (determinant +1) rotation, so it also captures
    axis reflections (sign flips) -- both are needed when matching
    permutation- or bootstrap-resampled canonical variates back to a
    reference fit, since resampling can induce either (Xia et al., 2018,
    *Nat. Commun.*; McIntosh & Lobaugh, 2004, *NeuroImage*).

    Args:
        reference: Array of shape (n, k).
        target: Array of shape (n, k), matched row-for-row with
            ``reference`` (e.g. the same features/variables in the same
            order), but not necessarily in the same column (component)
            order or sign.

    Returns:
        Orthogonal matrix of shape (k, k) such that ``target @ R`` is
        optimally aligned to ``reference`` in the least-squares sense.

    Raises:
        ValueError: If ``reference`` and ``target`` don't have the same
            shape.

    Examples:
        >>> import numpy as np
        >>> from cca_zoo.model_selection import procrustes_rotation
        >>> rng = np.random.default_rng(0)
        >>> reference = rng.standard_normal((20, 3))
        >>> true_rotation, _ = np.linalg.qr(rng.standard_normal((3, 3)))
        >>> target = reference @ true_rotation.T
        >>> recovered = procrustes_rotation(reference, target)
        >>> np.allclose(target @ recovered, reference, atol=1e-8)
        True
    """
    if reference.shape != target.shape:
        raise ValueError(
            "reference and target must have the same shape, got "
            f"{reference.shape} and {target.shape}."
        )
    m = target.T @ reference
    u, _, vt = np.linalg.svd(m)
    rotation: np.ndarray = u @ vt
    return rotation