Tag

Data Labeling

Data labeling is the process of tagging raw data (images, text, audio, etc.) to make it usable for training machine learning models. The quality, speed, and consistency of data labeling can significantly impact the value created by enterprise AI applications.

All articles on Data Labeling

Recap: The Future of Data-Centric AI Event
Main takeaways from The Future of Data-Centric AI Event We recently hosted The Future of Data-Centric AI, where academia, research, and industry experts and practitioners came together to discuss the shift from model-centric AI development to data-centric AI and what lies ahead. This post gives you a quick overview of the event and top takeaways from over eight hours of
October 11, 2021
Aarti Bagul
Web Virtualization — Optimizing Data-Intensive App Performance
Frontend Development Best Practices for Working With Lots of Data From Snorkel AI Engineering As a frontend engineer, it’s often easy to run into limitations when scaling large applications. At Snorkel AI, we often run into times where our users work with data that scales into the gigabytes when using Snorkel Flow. We have built Snorkel Flow around two core
September 16, 2021
Shubham Naik
Multi-Label Classification, Sequence Labeling, and More
Snorkel Flow LTS Release Summer ‘21 By adopting Snorkel Flow, a data-centric AI development platform powered by programmatic labeling, our customers have changed how they build and deploy AI applications. We’ve seen our customers save tens-of-millions of dollars in manual labeling costs and person-years of time by applying weak supervision with Snorkel Flow.Over the last few months, we’ve been hard
September 15, 2021
Patrick Kolencherry
Applying Weak Supervision Research
ScienceTalks with Paroma Varma In this episode of Science Talks, Snorkel AI’s Braden Hancock chats with Paroma Varma – a co-founder of Snorkel AI and one of the first and leading contributors to the Snorkel project. We discuss Paroma’s path into machine learning, her work in optimization and signal processing during her undergrad, weak supervision and image data during her
September 13, 2021
Team Snorkel
The Future of Data-Centric AI – Virtual Live Event
Join the live discussion. Learn how to unlock data-centric AI and make AI development practical in your organization Working with vast unstructured and unlabeled data is one of the bottlenecks in the machine learning lifecycle. Machine learning models can only get as reliable and accurate as the data being fed to them. With a data-centric approach 1, your data science
August 31, 2021
Team Snorkel
Snorkel AI Raises $85m Series C at $1b Valuation for Data-Centric AI
We started the Snorkel project at the Stanford AI lab in 2015 around two core hypotheses:
August 9, 2021
Alex Ratner
Developing and Managing Systems to Extract Structured Data
Machine Learning Whiteboard (MLW) Open-source Series Earlier this year, we started our machine learning whiteboard (MLW) series, an open-invite space to brainstorm ideas and discuss the latest papers, techniques, and workflows in the AI space. We emphasize an informal and open environment to everyone interested in learning about machine learning.In this episode, Manan Shah dives into “Glean: Structured Extractions from
August 2, 2021
Team Snorkel
Image
How to Use Snorkel to Build AI Applications
The how, what, and why of Snorkel’s programmatic data labeling approach and the state-of-the-art Snorkel Flow platform. The year was 2015. For the first time, machine learning (ML) had outperformed humans in the annual ImageNet challenge.
July 9, 2021
Braden Hancock
Multi-Resolution Weak Supervision for Sequential Data
Machine Learning Whiteboard (MLW) Open-source Series Our machine learning whiteboard (MLW) is an open-invite space to brainstorm ideas and discuss the latest papers, techniques, and workflows in the AI space. We emphasize an informal and open environment to everyone interested in discovering more about machine learning.In this episode, Hiromu Hota, Vincent Sunn Chen, Daniel Y. Fu, and Frederic Sala dive
June 25, 2021
Team Snorkel
Weak Supervision in Biomedicine
In this episode of Science Talks, Snorkel AI’s Braden Hancock chats with Jason Fries – a research scientist at Stanford University’s Biomedical Informatics Research lab and Snorkel Research, and one of the first contributors to the Snorkel open-source library. We discuss Jason’s path into machine learning, empowering doctors and scientists with weak supervision, and utilizing organizational resources in biomedical applications of Snorkel. This episode is part
June 16, 2021
Team Snorkel
Training Classifiers With Natural Language Explanations
Machine Learning Whiteboard (MLW) Open-source Series Earlier this year, we started our machine learning whiteboard (MLW) series, an open-invite space to brainstorm ideas and discuss the latest papers, techniques, and workflows in the AI space. We emphasize an informal and open environment to everyone interested in learning about machine learning.In this episode, our Co-founder and Head of Technology. Braden Hancock
May 24, 2021
Team Snorkel
Applying Information Theory to ML With Fred Sala
In this episode of Science Talks, Frederic Sala – an assistant professor of Computer Science at the University of Wisconsin Madison and a research scientist at Snorkel discusses his path into machine learning, the central thesis that ties together his multidisciplinary research, his thoughts on the future of weak supervision, as well as his decision to go into academia.
May 19, 2021
Team Snorkel
3 Impractical Assumptions About AI to Avoid
Impractical ML assumptions are made every day in research, which limit its adoption. In the real world, these assumptions do not hold up. Learn more about how to avoid making these assumptions about AI application development.
May 4, 2021
Braden Hancock
Introducing Application Studio and Announcing Our $35m Series B Funding
Over the past year, we’ve worked hard to deliver Snorkel Flow, the first AI platform to provide all the power of machine learning without the pains of hand-labeling. Snorkel Flow lets you label data programmatically, train models flexibly, improve performance iteratively, and deploy AI applications quickly. We are incredibly proud of the value that our customers, including two of the
April 5, 2021
Alex Ratner
Debugging AI Applications Pipeline
We’ll analyze major sources of errors during the four steps of building AI applications: data labeling, feature engineering, model training, and model evaluation.
February 3, 2021
Team Snorkel
Image
How To Overcome Practical Challenges for AI in the Public Sector
AI is already transforming the business of government. But the positive impacts of this transformation, from increasing the efficiency of public services to enhancing the effectiveness of tax dollars, are still in the earliest stages. Public sector organizations generally have access to the same talent, software models, and hardware infrastructure as any private sector company, but they face a number of relatively unique practical challenges that hinder their operationalization of AI.
January 7, 2021
Charlie Greenbacker
How To Overcome Practical Challenges for AI in Finance
Advancements in artificial intelligence promise efficiency gains for financial institutions. AI-powered applications can revolutionize an organization’s risk management, fraud detection, compliance monitoring, and other processes. Financial services companies have smart data scientists and good infrastructure needed for deploying AI. But their ability to rapidly develop and deploy AI applications is hampered by several unique challenges.
December 29, 2020
Manas Joglekar
Meet a Snorkeler at an Upcoming Event
We love meeting people in the data science and machine learning community. Here are a few upcoming events where you can meet Snorkelers.
November 17, 2020
Team Snorkel
How to Overcome Practical Challenges for AI in Healthcare
There’s a lot of excitement about the potential for AI to improve healthcare. This is driven by compelling advances across a wide range of applications including drug discovery, radiology, pathology, electronic medical record (EMR) intelligence, clinical trials, and more. There are also many challenges for development and deployment of AI for healthcare.
November 9, 2020
Brandon Yang
Image
Snorkel AI: Putting Data First in ML Development
Today I’m excited to announce Snorkel AI’s launch out of stealth! Snorkel AI, which spun out of the Stanford AI Lab in 2019, was founded on two simple premises: first, that the labeled training data machine learning models learn from is increasingly what determines the success or failure of AI applications. And second, that we can do much better than labeling this
July 14, 2020
Alex Ratner