Tag

Data Labeling

Data labeling is the process of tagging raw data (images, text, audio, etc.) to make it usable for training machine learning models. The quality, speed, and consistency of data labeling can significantly impact the value created by enterprise AI applications.

All articles on Data Labeling

How Georgetown University's CSET uses Snorkel Flow to build NLP applications to inform policy research banner
How Georgetown University’s CSET uses Snorkel Flow to build NLP applications to inform policy research
Georgetown University’s CSET is building next-generation NLP applications using Snorkel Flow to classify complex research documents. Snorkel Flow drastically reduced labeling, model training, and iteration time and better equipped CSET’s data science team to collaborate closely with analysts to gather, process, and interpret data at scale. 
December 19, 2022
Nick Harvey
Image
Snorkel AI Partners with Advanced Analytics Consultancy Aimpoint Digital
Snorkel AI is delighted to announce a partnership with Aimpoint Digital, a premier analytics firm specializing in AI application development that builds, operationalizes, and scales data science solutions for biopharma, manufacturing, retail, and other major industries. Aimpoint Digital leads the industry in solving complex challenges and exploiting value-generating opportunities for organizations of all sizes through data. The company helps clients
December 12, 2022
Friea Berg
Image
Supercharge data scientist and domain expert collaboration with Comments and Tags in Snorkel Flow
Labeling data manually can be a grind. Snorkel Flow slashes labeling time from months to minutes by allowing data scientists and domain experts collaborate through labeling functions. Snorkel Flow offers two unique capabilities that further supercharge that collaboration: Comments and Tags.
December 9, 2022
Marty Moesta
Image
Snorkel AI Team presents research at NeurIPS 2022
The Snorkel AI team will present five research papers advancing weak supervision and programmatic labeling at the NeurIPS 2022 conference that started this week.
November 29, 2022
Team Snorkel
Image
Deepening Snorkel AI’s partnership with Microsoft Azure AI
Snorkel AI is excited to build on our partnership with Microsoft Azure to help enterprises and government agencies solve their most impactful problems and unlock value from their data using AI. Learn how Azure customers can easily deploy Snorkel Flow on their Azure cloud infrastructure to accelerate AI application development with data-centric workflows and programmatic labeling.
November 22, 2022
Henry Ehrenberg
Image
Data-centric Foundation Model Development: Bridging the gap between foundation models and enterprise AI
Introducing new capabilities for Data-centric Foundation Model Development in Snorkel Flow Powerful new large language or foundation models (FMs) like GPT-3, Stable Diffusion, BERT, and more have taken the AI space by storm, going viral—even beyond technical practitioners—thanks to incredible capabilities around text generation, image synthesis, and more. However, enterprises face fundamental barriers to using these foundation models on real,
November 17, 2022
Alex Ratner
Image
Building an NLP application to analyze ESG factors in Earnings Calls using Snorkel Flow
Create a data-centric AI application using Snorkel Flow to save your analysts time of manual labeling and information extraction related to environmental, social, and governance (ESG) factors from earnings call transcripts. Rapidly and accurately extract all existing and new factors from the transcripts to make the right investment decision.
November 3, 2022
Amir Imani
Image
Building Trustworthy AI applications with data-centric AI
AI is generally accepted as necessary for organizations across private and public sectors to build (or maintain) a competitive advantage. However, a major challenge to adopting AI successfully is our ability to build reliable, predictable, and equitable solutions. A critical flaw with traditional approaches to developing AI is the reliance on hand-labeled training datasets and/or “pre-trained” black-box models that are effectively ungovernable and unauditable. In this article, we explore the motivations and challenges for Trustworthy AI that we’ve encountered and discuss how core tenants of Data-Centric AI, including programmatic labeling, help ameliorate them.
October 4, 2022
Arjun Prakash
Image
Top-10 US bank uses AI/ML to triage loan documents based on risk exposure
To meet the requirements of unexpected regulatory changes brought on by the pandemic, a top-10 US bank needed to urgently adapt its underperforming model-centric artificial intelligence and machine learning development approach to a data-centric one. The team used Snorkel Flow to automatically classify thousands of loan documents and extract critical clauses in just 24 hours, saving loan managers thousands of hours of manual document review.
September 30, 2022
Nick Harvey
Image
How Schlumberger uses Snorkel Flow to enhance proactive well management
Schlumberger is the world’s leading provider of technology and services for the energy industry, operating in over 120 countries. The company provides well maintenance and analytics services to the world’s biggest oil companies, and it believes that large-scale data analysis and artificial intelligence/machine learning will help them remain a leader in the market. One way they’ve been able to achieve this is by building their own AI application using Snorkel Flow to automatically extract geological entities and critical field data across a variety of document structures and report types they receive from their customers.
September 30, 2022
Nick Harvey
Image
Summer 2022 Snorkel Flow release roundup
On the heels of the second annual Future of Data-Centric AI event, we’re energized by what we learned from data scientists, machine learning engineers, and AI leaders who are adopting data-centric approaches to accelerate AI success. The Snorkel Flow platform provides these teams with a seamless workflow across training data creation, model training, and analysis—the scaffolding to make data-centric AI
August 30, 2022
Molly Friederich
Image
Introducing Continuous Model Feedback to drive rapid data quality improvement
Continuous Model Feedback, available in beta as part of the new Studio experience, is Snorkel Flow’s latest capabilities to make training data creation and model development more integrated, automated, and guided.
August 29, 2022
Molly Friederich
Image
The Future of Data-Centric AI 2022 day 2 highlights
Snorkel AI just hosted the second day of The Future of Data-Centric AI conference 2022. Across 40+ sessions, 50+ Data scientists, ML engineers, and AI leaders came together to share insights, best practices, and research on adopting data-centric approaches with thousands of attendees from all around the world. Aarti Bagul, a Snorkel AI ML Solutions Engineer and one of the
August 5, 2022
Louis Bouchard
Image
The Future of Data-Centric AI 2022 day 1 highlights
Snorkel AI just hosted the first day of The Future of Data-Centric AI conference 2022. This conference brings together data scientists, ML engineers, and AI leaders to share insights, best practices, and research on how to evolve the ML lifecycle from model-centric to data-centric approaches. This conference takes place over two days with 40+ sessions, 50+ speakers, and thousands of
August 4, 2022
Louis Bouchard
Information extraction case studies for 10-Ks
10-Ks information extraction case studies
Building NLP techniques to understand 10-Ks is time-consuming, costly, and challenging. In this post, Machine Learning Engineer, Aarti Bagul discusses three information extraction case studies on how banks around the world are building highly accurate NLP applications using Snorkel Flow’s AI platform. From retail banking to hedge fund investing, NLP is used across the financial industry. By processing and extracting
July 6, 2022
Team Snorkel
Image
Introducing Cluster View: Instant data insight made actionable to speed AI development
Programmatic labeling moves a classic technique from interesting to high-impact So much of real-world AI development entails working with text data that’s messy — in fact, 80%+ of enterprise data is unstructured. And while state-of-the-art models get a lot of the glory, creating the training data that conveys what your model needs to learn is more often the biggest determiner of AI
June 30, 2022
Molly Friederich
Image
Data-centric approaches to multi-label classification
AI systems are well-suited to tasks involving recognizing and predicting data patterns. Supervised classification systems categorize unseen data into a finite set of discrete classes by learning from millions of hand-labeled labeled sample points. These classifiers are powerful business tools – they automate document sorting, customer sentiment analysis, sales performance, and other distinct business problems. However, they also require an
June 29, 2022
Kanyes Thaker
Guidelines and best practices for annotation of data
Data annotation guidelines and best practices
What is data annotation? Data annotation refers to the process of categorizing and labeling data for training datasets. This process plays a critical role in preparing data for machine learning models, as high-quality training data enables more accurate predictions and insights. In order for a training dataset to be usable, it must be categorized appropriately and annotated for a specific
June 28, 2022
Anastassia Kornilova
Image
3 ways to use Snorkel’s Labeling Functions
Labeling functions are fundamental building blocks of programmatic labeling that encode diverse sources of weak labeling signals to produce high-quality labeled data at scale. Let’s start with the core motivation for labeling functions: over time, every major commercial organization and government agency builds various valuable, often bespoke knowledge resources. These resources include employee expertise, wikis and ontologies, business logic, and
June 24, 2022
Nic Acton
Image
Clinical entity classification in electronic health records
Research recap: Ontology-driven weak supervision for clinical entity classification in electronic health records (EHRs)  In this post, I have summarized the research published in this academic paper, Ontology-driven weak supervision for clinical entity classification in electronic health records by Jason Fries et al. This paper was published in Nature Communications in 2021.Problem statement Electronic health records (EHR) contain a rich
June 17, 2022
Nazanin Makkinejad
Image
Building AI models for financial document processing best practices
Highlighting the best practices for building and deploying AI models for financial document processing applications AI has massive potential in the financial industry. Building AI models to automate information extraction, fraud detection, and compliance monitoring can provide efficient and faster responses and support repurposing domain experts’ labor to more meaningful tasks. Developing AI models is not just about having models
June 15, 2022
Hoang Tran
Trustworthy AI, image by Tara Winstead
The benefits of programmatic labeling for trustworthy AI
The following post is based on a talk discussing the benefits of programmatic labeling for trustworthy AI, which was presented as part of the Trustworthy AI: A Practical Roadmap for Government event that took place this past April, with Snorkel AI Co-founder and Head of Technology, Braden Hancock. If you would like to watch Braden’s presentation, we have included it
June 9, 2022
Team Snorkel
Image
Named entity extraction and recognition with Snorkel Flow
If you were ever amazed at how Google accurately finds the answer to your question just by a few keywords, you’ve witnessed the power of named entity recognition (NER). By quickly and accurately identifying different entities in a sea of unstructured articles, like names of people, places, and organizations, the search engine can figure out each article’s main topics and
June 7, 2022
April Guo
James Zou portrayed
A data-centric perspective on trustworthy and interpretable AI
The future of data-centric AI talk series In this talk, Assistant Professor of Biomedical Data Science at Stanford University, James Zou, discusses the work he and his team have been doing from a data-centric perspective to trustworthy and interpretable AI. If you would like to watch James’ presentation, we have included it below, or you can find the entire event
June 6, 2022
Team Snorkel
Image
Government keynote presentation by FBI CTO Gregory Ihrie
Gregory Ihrie is the Chief Technology Officer for the FBI, responsible for technology, innovation, and strategy. He also leads the FBI’s efforts in advancing the bureau’s management, policy, and governance of AI systems. Ihrie chairs the FBI’s Scientific Working Group on Artificial Intelligence, as well as the Department of Justice’s AI Committee of Interest. He is one of three officers
June 4, 2022
Team Snorkel
The future of data-centric AI presented by Snorkel AI
What to expect at The Future of Data-Centric AI 2022
30+ sessions by 40+ speakers in 2 action-packed days Last year we organized The Future of Data-Centric AI conference to explore the shift from model-centric to data-centric AI. Speakers included researchers and industry experts such as Andrew Ng (Landing AI), Anima Anandkumar (NVIDIA), Chris Re (Stanford AI Lab), Michael DAndrea (Genentech), Skip McCormick (BNY Mellon), Imen Grida Ben Yahia (Orange)
June 1, 2022
Devang Sachdev
Image
Auto LF generation: Lots of little models, big benefits
Constructing labeling functions (LFs) is at the heart of using weak supervision. We often think of these labeling functions as programmatic expressions of domain expertise or heuristics. Indeed, much of the advantage of weak supervision is that we can save time—writing labeling functions and applying them to data at scale is much more efficient compared to hand-labeling huge numbers of
May 31, 2022
Fred Sala
Image
Building a COVID fact-checking system with external knowledge
Powerful resources to leverage as labeling functions In this post, we’ll use the COVID-FACT dataset to demonstrate how to use existing resources as labeling functions (LFs), to build a fact-checking system. The COVID-FACT dataset contains 4086 claims about the COVID-19 pandemic; it contains claims, evidence for the claims, and contradictory claims refuted by the evidence. The evidence retrieval is formulated
May 26, 2022
Annie Yang
Image
Snorkel AI FAQ
Browse through these FAQ to find answers to commonly raised questions about Snorkel AI, Snorkel Flow, and data-centric AI development. Have more questions? Contact us. Programmatic labeling Use cases 1. What is a labeling function? A Labeling Function (LF) is an arbitrary function that takes in a data point and outputs a proposed label or abstains. The logic used to
May 25, 2022
Team Snorkel
Image
Panel discussion: Academic and industry perspectives on ethical AI
This post showcases a panel discussion on the academic and industry perspectives of ethical AI, which was moderated by Director of Federal Strategy and Growth, Alexis Zumwalt, Fouts Family Early Career Professor and Lead of Ethical AI (NSF AI Institute AI4OPT), Georgia Institute of Technology, Swati Gupta, Chief Data Officer, Department of the Navy, Thomas Sasalsa, Senior Manager of Responsible
May 24, 2022
Team Snorkel