Sketching Datasets for Large-Scale Learning (long version)

This article considers "sketched learning," or "compressive learning," an approach to large-scale machine learning where datasets are massively compressed before learning (e.g., clustering, classification, or regression) is performed. In particular, a "sketch" is first constructed by computing carefully chosen nonlinear random features (e.g., random Fourier features) and averaging them over the whole dataset. Parameters are then learned from the sketch, without access to the original dataset. This article surveys the current state-of-the-art in sketched learning, including the main concepts and algorithms, their connections with established signal-processing methods, existing theoretical guarantees-on both information preservation and privacy preservation, and important open problems.

Domaines

Traitement du signal et de l'image [eess.SP] Apprentissage [cs.LG]

Fichier principal

SPM_paper.pdf (7.94 Mo)

SPM_paper.bcf (108.64 Ko)

SPM_paper.run.xml (2.36 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Rémi Gribonval : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-02909766

Soumis le : mardi 26 janvier 2021-17:42:55

Dernière modification le : jeudi 4 avril 2024-18:17:55

Dates et versions

hal-02909766 , version 1 (04-08-2020)

hal-02909766 , version 2 (26-01-2021)

Identifiants

HAL Id : hal-02909766 , version 2
ARXIV : 2008.01839

Citer

Rémi Gribonval, Antoine Chatalic, Nicolas Keriven, Vincent Schellekens, Laurent Jacques, et al.. Sketching Datasets for Large-Scale Learning (long version). 2021. ⟨hal-02909766v2⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

ENS-LYON UNIV-LYON3 UNIV-RENNES1 UGA CNRS INRIA UNIV-LYON1 UNIV-LYON2 INSA-LYON INSA-RENNES IRISA GIPSA CENTRALESUPELEC INRIA2 UR1-MATH-STIC UR1-UFR-ISTIC UNIV-RENNES INSA-GROUPE UDL GIPSA-GAIA ANR UR1-MATH-NUM

308 Consultations

157 Téléchargements