Column Names as Contracts - Inspired by dplyr, Available in dbt by Emily Riederer
Visit https://rstats.ai/nyr to learn more.
Abstract: dplyr’s select helpers exemplify how the tidyverse uses opinionated design to push users into the pit of success. The ability to efficiently operate on names incentivizes good naming patterns and creates efficiency in data wrangling and validation.
However, in a polyglot world, users may find they must leave the pit when comparable syntactic sugar is not accessible in other languages like python and SQL.
In this talk, I will explain how dplyr’s select helpers inspired my approach to ‘column name contracts’, how good naming systems can help supercharge data management with packages like {dplyr} and {pointblank}, and my experience building the {dbtplyr} to port this functionality to dbt for building complex SQL-based data pipelines.
Bio: Emily is a Senior Manager at Capital One where she has led a variety of teams focused on actionable analytics, intuitive data product development, and fit-for-purpose data science solutions. Outside of work, she is an active member of the #rstats and broader data communities. She enjoys doing pro-bono data volunteer work, writing about the data space (with her work appearing in the R Markdown Cookbook, 97 Things Data Engineers Should Know, and on her website), and supporting open-source software and open-knowledge sharing with her roles as on the editorial board for rOpenSci and a frequent technical reviewer for CRC Press.
Twitter: / emilyriederer
Presented at the 2023 New York R Conference (July 14, 2023)