Johannes Hötter is a data scientist and engineer, and the co-founder of kern. He is passionate about machine learning and data management, and especially about the things people build when the two come together.
At kern, Johannes focuses on one of the least glamorous but most decisive parts of applied NLP: the data. Data scientists spend an outsized share of their time labeling and managing data, and kern’s open-source tools such as Refinery and Bricks are built to take that pain away. Refinery brings structure to exploratory data labeling and management at large scale, while Bricks provides a library of labeling heuristics, including GPT-driven rules, that can be combined with active learning and crowd labels into ensembles of virtual annotators. The result is that teams can build high-quality training datasets without hand-labeling every example.
Through his work, Johannes advocates a pragmatic approach to modern NLP: combine weak supervision with large language models and solid data management foundations such as embeddings and Hugging Face models, and make deliberate choices about open source versus commercial productization so that tools actually fit the workflows of engineering teams.
Johannes Hötter