Tag

Data Development

Data development encompasses the processes of curating, organizing, and preparing datasets for use in machine learning and AI projects. This includes data sourcing, cleaning, labeling, and augmenting, ensuring that the data used is high-quality and relevant. Data-centric approaches prioritize the value of data itself, often leading to more reliable and efficient model outcomes.

All articles on Data Development

Building a Successful AI Startup
ScienceTalks with Saam Motamedi We at Snorkel AI have received many requests from data scientists and machine learning engineers who aspire to be founders, where do they start and how should they get started on their entrepreneurial journey? We genuinely believe that data scientists and machine learning engineers will build the next generation of mega-enterprises. Over the summer, we’ve recorded
October 18, 2021
Team Snorkel
Recap: The Future of Data-Centric AI Event
Main takeaways from The Future of Data-Centric AI Event We recently hosted The Future of Data-Centric AI, where academia, research, and industry experts and practitioners came together to discuss the shift from model-centric AI development to data-centric AI and what lies ahead. This post gives you a quick overview of the event and top takeaways from over eight hours of
October 11, 2021
Aarti Bagul
Web Virtualization — Optimizing Data-Intensive App Performance
Frontend Development Best Practices for Working With Lots of Data From Snorkel AI Engineering As a frontend engineer, it’s often easy to run into limitations when scaling large applications. At Snorkel AI, we often run into times where our users work with data that scales into the gigabytes when using Snorkel Flow. We have built Snorkel Flow around two core
September 16, 2021
Shubham Naik
Sliceline: Fast, Linear-Algebra-Based Slice Finding for ML Model Debugging
Diving Into SliceLine – Machine Learning Whiteboard (MLW) Open-source Series Earlier this year, we started our machine learning whiteboard (MLW) series, an open-invite space to brainstorm ideas and discuss the latest papers, techniques, and workflows in the AI space. We emphasize an informal and open environment to everyone interested in learning about machine learning.In this episode, Kaushik Shivakumar dives into
September 8, 2021
Team Snorkel
The Future of Data-Centric AI – Virtual Live Event
Join the live discussion. Learn how to unlock data-centric AI and make AI development practical in your organization Working with vast unstructured and unlabeled data is one of the bottlenecks in the machine learning lifecycle. Machine learning models can only get as reliable and accurate as the data being fed to them. With a data-centric approach 1, your data science
August 31, 2021
Team Snorkel
Snorkel AI Raises $85m Series C at $1b Valuation for Data-Centric AI
We started the Snorkel project at the Stanford AI lab in 2015 around two core hypotheses:
August 9, 2021
Alex Ratner
Developing and Managing Systems to Extract Structured Data
Machine Learning Whiteboard (MLW) Open-source Series Earlier this year, we started our machine learning whiteboard (MLW) series, an open-invite space to brainstorm ideas and discuss the latest papers, techniques, and workflows in the AI space. We emphasize an informal and open environment to everyone interested in learning about machine learning.In this episode, Manan Shah dives into “Glean: Structured Extractions from
August 2, 2021
Team Snorkel
Weak Supervision in Biomedicine
In this episode of Science Talks, Snorkel AI’s Braden Hancock chats with Jason Fries – a research scientist at Stanford University’s Biomedical Informatics Research lab and Snorkel Research, and one of the first contributors to the Snorkel open-source library. We discuss Jason’s path into machine learning, empowering doctors and scientists with weak supervision, and utilizing organizational resources in biomedical applications of Snorkel. This episode is part
June 16, 2021
Team Snorkel
Applying Information Theory to ML With Fred Sala
In this episode of Science Talks, Frederic Sala – an assistant professor of Computer Science at the University of Wisconsin Madison and a research scientist at Snorkel discusses his path into machine learning, the central thesis that ties together his multidisciplinary research, his thoughts on the future of weak supervision, as well as his decision to go into academia.
May 19, 2021
Team Snorkel
3 Impractical Assumptions About AI to Avoid
Impractical ML assumptions are made every day in research, which limit its adoption. In the real world, these assumptions do not hold up. Learn more about how to avoid making these assumptions about AI application development.
May 4, 2021
Braden Hancock
Building Industrial-Strength NLP Applications With Ines Montani
In this episode of Science Talks, Explosion AI’s Ines Montani sat down with Snorkel AI’s Braden Hancock to discuss her path into machine learning, key design decisions behind the popular spaCy library for industrial-strength NLP, the importance of bringing together different stakeholders in the ML development process, and more.This episode is part of the #ScienceTalks video series hosted by the Snorkel AI team. You
April 29, 2021
Team Snorkel
Introducing Application Studio and Announcing Our $35m Series B Funding
Over the past year, we’ve worked hard to deliver Snorkel Flow, the first AI platform to provide all the power of machine learning without the pains of hand-labeling. Snorkel Flow lets you label data programmatically, train models flexibly, improve performance iteratively, and deploy AI applications quickly. We are incredibly proud of the value that our customers, including two of the
April 5, 2021
Alex Ratner
Measuring NLP Progress With Sebastian Ruder
In this episode of Science Talks, Sebastian Ruder, Research Scientist at DeepMind, shares his thoughts on making AI practical with Snorkel AI’s Braden Hancock. This conversation covers progress made in the NLP domain with emerging research, new benchmarks like SuperGLUE, rich repositories and news sources that keep you in the loop and on top of what’s new in NLP, and more.
March 10, 2021
Team Snorkel
Machine Learning Production Myths
Takeaways from MLSys Seminars with Chip HuyenIn November, I had the opportunity to come back to Stanford to participate in MLSys Seminars, a series about Machine Learning Systems. It was great to see the growing interest of the academic community in building practical AI applications. Here is a recording of the talk.The talk was originally about the principles of good
December 23, 2020
Chip Huyen
How to Overcome Practical Challenges for AI in Healthcare
There’s a lot of excitement about the potential for AI to improve healthcare. This is driven by compelling advances across a wide range of applications including drug discovery, radiology, pathology, electronic medical record (EMR) intelligence, clinical trials, and more. There are also many challenges for development and deployment of AI for healthcare.
November 9, 2020
Brandon Yang
Image
Snorkel AI: Putting Data First in ML Development
Today I’m excited to announce Snorkel AI’s launch out of stealth! Snorkel AI, which spun out of the Stanford AI Lab in 2019, was founded on two simple premises: first, that the labeled training data machine learning models learn from is increasingly what determines the success or failure of AI applications. And second, that we can do much better than labeling this
July 14, 2020
Alex Ratner