Tag

Annotation

Annotation involves labeling raw data—such as text, images, or video—with relevant tags to prepare it for machine learning models. Accurate data annotation forms the foundation for training AI systems. Models learn from patterns in labeled data and then mimic them.

Whether manual or automated, annotation must be consistent and scalable, especially when handling large datasets.

All articles on Annotation

Image
Comcast’s data-centric approach to speech interfaces
Jan Neumann, Vice President of Machine Learning for Comcast Applied AI and Discovery, describes Comcast’s data-centric AI approach to speech.
February 14, 2023
Team Snorkel
Image
Snorkel AI and Google Cloud accelerate AI innovation
Snorkel AI is teaming up with Google Cloud to help F500 companies and AI innovators solve their most difficult problems.
February 2, 2023
Friea Berg
Liger: Fusing foundation model embeddings & weak supervision blog image
How Foundation Models bolster programmatic labeling
Snorkel CEO Alex Ratner interviews Mayee Chen about how Liger improves the effectiveness of programmatic labeling through foundation model embeddings.
January 26, 2023
Team Snorkel
Image
Prompting and weak supervision to build better, smaller models
Snorkel AI co-founder and CEO Alex Ratner recently interviewed several Snorkel researchers about their published academic papers. In this video, Alex talks with Ryan Smith, Senior Applied Scientist at Snorkel, about the work he did on using foundation models to build compact, deployable, and effective models.
January 19, 2023
Team Snorkel
Image
Combining human and artificial intelligence with human-in-the-loop ML | FDCAI
More components in an ML lifecycle are designed to run on autopilot, but some tasks require human-in-the-loop ML, an active research topic that has seen an increasing number of publications in the last 10 years.
December 28, 2022
Team Snorkel
How a top 3 US bank used Snorkel Flow to automate 10-K review for their analysts Banner image
How a top 3 US bank used Snorkel Flow to automate 10-K review for their analysts
A central innovation team at a top US bank wanted to modernize its AI development and data annotation processes in order to create a custom natural language processing (NLP) model that could extract important financial information from 10-Ks. Manually reviewing these documents was taking up valuable time that could be better spent assisting customers. The team used Snorkel Flow’s data-centric AI development process and programmatic labeling to train a customized NLP model that could accurately extract information on interest rate swaps.
December 23, 2022
Nick Harvey
How programmatic labeling can minimize data exposure blog banner
How programmatic labeling can minimize data exposure
MIT’s Technology Review reported this week that workers in Venezuela contracted by outsourced data annotation services provider shared customer data—low-angled pictures intended to be labeled, including one that featured a woman in a private moment in the bathroom—with each other on social media. Programmatic labeling could have minimized this.
December 21, 2022
Devang Sachdev
Image
Supercharge data scientist and domain expert collaboration with Comments and Tags in Snorkel Flow
Labeling data manually can be a grind. Snorkel Flow slashes labeling time from months to minutes by allowing data scientists and domain experts collaborate through labeling functions. Snorkel Flow offers two unique capabilities that further supercharge that collaboration: Comments and Tags.
December 9, 2022
Marty Moesta
Image
What can Data-Centric AI learn from data & ML engineering?
Databricks’ Chief Technologist: Data-Centric AI can learn from Data Engineering and ML Engineering in five ways: continuous updates, versioning, code-centric deployment, data privatization and actionable monitoring.
November 5, 2022
Team Snorkel
Image
Building an NLP application to analyze ESG factors in Earnings Calls using Snorkel Flow
Create a data-centric AI application using Snorkel Flow to save your analysts time of manual labeling and information extraction related to environmental, social, and governance (ESG) factors from earnings call transcripts. Rapidly and accurately extract all existing and new factors from the transcripts to make the right investment decision.
November 3, 2022
Amir Imani
Image
Summer 2022 Snorkel Flow release roundup
On the heels of the second annual Future of Data-Centric AI event, we’re energized by what we learned from data scientists, machine learning engineers, and AI leaders who are adopting data-centric approaches to accelerate AI success. The Snorkel Flow platform provides these teams with a seamless workflow across training data creation, model training, and analysis—the scaffolding to make data-centric AI
August 30, 2022
Molly Friederich
Guidelines and best practices for annotation of data
Data annotation guidelines and best practices
What is data annotation? Data annotation refers to the process of categorizing and labeling data for training datasets. This process plays a critical role in preparing data for machine learning models, as high-quality training data enables more accurate predictions and insights. In order for a training dataset to be usable, it must be categorized appropriately and annotated for a specific
June 28, 2022
Anastassia Kornilova
Image
Clinical entity classification in electronic health records
Research recap: Ontology-driven weak supervision for clinical entity classification in electronic health records (EHRs)  In this post, I have summarized the research published in this academic paper, Ontology-driven weak supervision for clinical entity classification in electronic health records by Jason Fries et al. This paper was published in Nature Communications in 2021.Problem statement Electronic health records (EHR) contain a rich
June 17, 2022
Nazanin Makkinejad
Trustworthy AI, image by Tara Winstead
The benefits of programmatic labeling for trustworthy AI
The following post is based on a talk discussing the benefits of programmatic labeling for trustworthy AI, which was presented as part of the Trustworthy AI: A Practical Roadmap for Government event that took place this past April, with Snorkel AI Co-founder and Head of Technology, Braden Hancock. If you would like to watch Braden’s presentation, we have included it
June 9, 2022
Team Snorkel
Image
Government keynote presentation by FBI CTO Gregory Ihrie
Gregory Ihrie is the Chief Technology Officer for the FBI, responsible for technology, innovation, and strategy. He also leads the FBI’s efforts in advancing the bureau’s management, policy, and governance of AI systems. Ihrie chairs the FBI’s Scientific Working Group on Artificial Intelligence, as well as the Department of Justice’s AI Committee of Interest. He is one of three officers
June 4, 2022
Team Snorkel
Image
Auto LF generation: Lots of little models, big benefits
Constructing labeling functions (LFs) is at the heart of using weak supervision. We often think of these labeling functions as programmatic expressions of domain expertise or heuristics. Indeed, much of the advantage of weak supervision is that we can save time—writing labeling functions and applying them to data at scale is much more efficient compared to hand-labeling huge numbers of
May 31, 2022
Fred Sala
Image
Snorkel AI FAQ
Browse through these FAQ to find answers to commonly raised questions about Snorkel AI, Snorkel Flow, and data-centric AI development. Have more questions? Contact us. Programmatic labeling Use cases 1. What is a labeling function? A Labeling Function (LF) is an arbitrary function that takes in a data point and outputs a proposed label or abstains. The logic used to
May 25, 2022
Team Snorkel
Image
Panel discussion: Academic and industry perspectives on ethical AI
This post showcases a panel discussion on the academic and industry perspectives of ethical AI, which was moderated by Director of Federal Strategy and Growth, Alexis Zumwalt, Fouts Family Early Career Professor and Lead of Ethical AI (NSF AI Institute AI4OPT), Georgia Institute of Technology, Swati Gupta, Chief Data Officer, Department of the Navy, Thomas Sasalsa, Senior Manager of Responsible
May 24, 2022
Team Snorkel
Image
Event recap: Adopting trustworthy AI for government
We’re currently experiencing such a rapid AI revolution and adoption of technologies, ranging from autonomous cars to virtual assistants and robotic surgeries and so much more, making it challenging for our government agencies to keep up. Especially when adding AI technologies to the mix, it can be even harder to manage.The crucial adoption of trustworthy AI and its successful integration
May 23, 2022
Alexis Zumwalt
Liger: Fusing foundation model embeddings & weak supervision
Showcasing Liger—a combination of foundation model embeddings to improve weak supervision techniques. Machine learning whiteboard (MLW) open-source series In this talk, Mayee Chen, a PhD student in Computer Science at Stanford University focuses on her work combining weak supervision and foundation model embeddings that improve two essential aspects of current weak supervision techniques. Check out the full episode here or
May 9, 2022
Team Snorkel
Active learning: an overview
A primer on active learning presented by Josh McGrath. Machine learning whiteboard (MLW) open-source series This video defines active learning, explores variants and design decisions made within active learning pipelines, and compares it to related methods. It contains references to some seminal papers in machine learning that we find instructive. Check out the full video below or on Youtube. Additionally, a
May 4, 2022
Josh McGrath
Image
Bill of materials for responsible AI: collaborative labeling
In our previous posts, we discussed how explainable AI is crucial to ensure the transparency and auditability of your AI deployments and how trustworthy AI adoption and its successful integration into our country’s critical infrastructure and systems are paramount. In this post, we dive into making trustworthy and responsible AI possible with Snorkel Flow, the data-centric AI platform for government and federal agencies. Collaborative labeling and
April 28, 2022
Alexis Zumwalt
Image
How to better govern ML models? Hint: auditable training data
ML models will always have some level of bias. Rather than relying on black-box algorithms, how can we make the entire AI development workflow more auditable? How do we build applications where bias can be easily detected and quickly managed? Today, most organizations focus their model governance efforts on investigating model performance and the bias within the predictions. Data science
April 6, 2022
Jonathan Dahlberg
Image
Algorithms that leverage data from other tasks with Chelsea Finn
The Future of Data-Centric AI Talk Series Background Chelsea Finn is an assistant professor of computer science and electrical engineering at Stanford University, whose research has been widely recognized, including in the New York Times and MIT Technology Review. In this talk, Chelsea talks about algorithms that use data from tasks you are interested in and data from other tasks.
March 31, 2022
Team Snorkel
Image
Weak Supervision Modeling with Fred Sala
Understanding the label model. Machine learning whiteboard (MLW) open-source series Background Frederic Sala, is an assistant professor at the University of Wisconsin-Madison, and a research scientist at Snorkel AI. Previously, he was a postdoc in Chris Re’s lab at Stanford. His research focuses on data-driven systems and weak supervision. In this talk, Fred focuses on weak supervision modeling. This machine
March 17, 2022
Team Snorkel
Image
How AI can be used to rapidly respond to information warfare in the Russia-Ukraine conflict
Proliferating web technology has contributed to information warfare in recent conflicts. Artificial Intelligence (AI) can play a significant role in stemming disinformation campaigns, cyber-attacks, and informing diplomacy in the rapidly evolving situation in Ukraine. Snorkel AI is dedicated to supporting the National Security community and other enterprise organizations with state-of-the-art AI technology. We see this as our responsibility in the
February 28, 2022
Nic Acton
,
Charlie Greenbacker
How Genentech extracted information for clinical trial analytics with Snorkel Flow
Genentech, a global biotech leader and member of the Roche Group, leveraged Snorkel Flow to extract critical information from lengthy clinical trial protocol (CTP) pdf documents. They built AI applications that used NER, entity linking, text extraction, and classification models to determine inclusion/ exclusion criteria and to analyze Schedules of Assessments. Genentech’s team achieved 95-99% model accuracy by using Snorkel
February 26, 2022
Team Snorkel
Augmenting the clinical trial design process with information extraction
The future of data-centric AI talk series Background Michael DAndrea is the Principal Data Scientist at Genentech. He earned his MBA from Cornell University and a Master’s degree in Computing and Education from Columbia University. He currently works on using unstructured data sources for clinical trial analytics and his team is partnered with the Stanford “AI For Health” initiative as
February 22, 2022
Team Snorkel
Q4 LTS Release of Snorkel Flow
We’re excited to announce the Q4 2021 LTS release of Snorkel Flow, our data-centric AI development platform powered by programmatic labeling. This latest release introduces a number of new product capabilities and enhancements, from a streamlined programmatic data development interface, to enhanced auto-suggest for labeling functions, to new machine learning capabilities like AutoML, to significant performance enhancements for PDF data
February 8, 2022
Henry Ehrenberg
Image
The Principles of Data-Centric AI Development
The Future of Data-Centric AI Talk Series Background Alex Ratner is CEO and co-founder of Snorkel AI and an Assistant Professor of Computer Science at the University of Washington. He recently joined the Future of Data-Centric AI event, where he presented the principles of data-centric AI and where it’s headed. If you would like to watch his presentation in full,
January 25, 2022
Team Snorkel