Speaker: Jonas Mueller: Chief Scientist, Cleanlab
In applied ML projects, experienced data scientists know that improving data brings higher ROI than tinkering with models. However the process of finding and fixing problems in a dataset is highly manual (ad hoc ideas explored in Jupyter notebooks). Cleanlab develops open-source software to help make this process more: efficient (via novel algorithms that automatically detect certain issues in data) and systematic (with better coverage to detect different types of issues).
This talk will describe how high-level ideas from data-centric AI can be operationalized across a wide variety of datasets (image, text, tabular, etc). I will introduce novel algorithmic strategies to automatically identify various issues in data that we have researched and published papers on with extensive benchmarks. These include detection of label errors, bad data annotators, out-of-distribution examples, and other dataset problems that once identified can be easily addressed to significantly improve trained models. Thousands of data scientists have started using this sort of data-centric AI software, and results from a few case studies will be presented. I will conclude with a discussion of where the data-centric AI movement is headed next, and key obstacles that deserve more attention.