pyDVL is a library of stable implementations of algorithms for data valuation and influence function computation
-
Updated
Apr 13, 2026 - Python
pyDVL is a library of stable implementations of algorithms for data valuation and influence function computation
Code for ACL 2025 Main paper "Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning".
Data-efficient Fine-tuning for LLM-based Recommendation (SIGIR'24)
Learning Large-scale Neural Fields via Context Pruned Meta-Learning (NeurIPS 2023)
Jaehyung Kim et al's ACL 2023 paper on "infoVerse: A Universal Framework for Dataset Characterization with Multidimensional Meta-information"
Official repository of the paper "DiffProb: Data Pruning for Face Recognition" (accepted at FG 2025)
DataCull is a modular, light-weight data pruning library containing many dataset pruning (coreset selection) algorithm including the official Implementation of the paper, titled, RCAP: Robust, Class-Aware, Probab ilistic Dynamic Dataset Pruning
🌿 dPrune: A Framework for Data Pruning
code for the paper Beyond Neural scaling laws for fast proven robust certification of nearest prototype classifiers
Ask your data questions in plain English and get instant AI-powered insights, charts, and visualizations—no SQL or code required.
Add a description, image, and links to the data-pruning topic page so that developers can more easily learn about it.
To associate your repository with the data-pruning topic, visit your repo's landing page and select "manage topics."