Erik Altman | Synthetic Data Sets: Use Cases for the Financial Industry | Stanford HAI
Stanford
University
  • Stanford Home
  • Maps & Directions
  • Search Stanford
  • Emergency Info
  • Terms of Use
  • Privacy
  • Copyright
  • Trademarks
  • Non-Discrimination
  • Accessibility
© Stanford University.  Stanford, California 94305.
Skip to content
  • About

    • About
    • People
    • Get Involved with HAI
    • Support HAI
    • Subscribe to Email
  • Research

    • Research
    • Fellowship Programs
    • Grants
    • Student Affinity Groups
    • Centers & Labs
    • Research Publications
    • Research Partners
  • Education

    • Education
    • Executive and Professional Education
    • Government and Policymakers
    • K-12
    • Stanford Students
  • Policy

    • Policy
    • Policy Publications
    • Policymaker Education
    • Student Opportunities
  • AI Index

    • AI Index
    • AI Index Report
    • Global Vibrancy Tool
    • People
  • News
  • Events
  • Industry
  • Centers & Labs
Navigate
  • About
  • Events
  • Careers
  • Search
Participate
  • Get Involved
  • Support HAI
  • Contact Us

Stay Up To Date

Get the latest news, advances in research, policy work, and education program updates from HAI in your inbox weekly.

Sign Up For Latest News

Your browser does not support the video tag.
eventSeminar

Erik Altman | Synthetic Data Sets: Use Cases for the Financial Industry

Status
Past
Date
Wednesday, May 07, 2025 12:00 PM - 1:15 PM PST/PDT
Location
Gates Computer Science Building, Room 119 353 Jane Stanford Way Stanford, CA 94305
Topics
Finance, Business

IBM Synthetic Data Sets (SDS) have been created for use cases in the financial industry.  

Share
Link copied to clipboard!
Event Contact
Annie Benisch
abenisch@stanford.edu
2099183302

Related Events

Caroline Meinhardt, Thomas Mullaney, Juan N. Pava, and Diyi Yang | How Can AI Support Language Digitization and Digital Inclusion?
SeminarApr 15, 202612:00 PM - 1:15 PM
April
15
2026

What does digital inclusion look like in the age of AI? Over 6,000 of the world’s 7,000-plus living languages remain digitally disadvantaged.

Seminar

Caroline Meinhardt, Thomas Mullaney, Juan N. Pava, and Diyi Yang | How Can AI Support Language Digitization and Digital Inclusion?

Apr 15, 202612:00 PM - 1:15 PM

What does digital inclusion look like in the age of AI? Over 6,000 of the world’s 7,000-plus living languages remain digitally disadvantaged.

AI+Science: Accelerating Discovery
ConferenceMay 05, 20268:30 AM - 5:00 PM
May
05
2026

AI+Science: Accelerating Discovery is an interdisciplinary conference bringing together researchers across physics, mathematics, chemistry, biology, neuroscience, and more to examine how AI is reshaping scientific discovery. Experts will separate hype from reality, spotlighting where AI is already enabling genuine breakthroughs and where its limits and risks remain.

Conference

AI+Science: Accelerating Discovery

May 05, 20268:30 AM - 5:00 PM

AI+Science: Accelerating Discovery is an interdisciplinary conference bringing together researchers across physics, mathematics, chemistry, biology, neuroscience, and more to examine how AI is reshaping scientific discovery. Experts will separate hype from reality, spotlighting where AI is already enabling genuine breakthroughs and where its limits and risks remain.

Wolfgang Lehrach | Code World Models for General Game Playing
SeminarMay 13, 202612:00 PM - 1:15 PM
May
13
2026

While Large Language Models (LLMs) show promise in many domains, relying on them for direct policy generation in games often results in illegal moves and poor strategic play.

Seminar

Wolfgang Lehrach | Code World Models for General Game Playing

May 13, 202612:00 PM - 1:15 PM

While Large Language Models (LLMs) show promise in many domains, relying on them for direct policy generation in games often results in illegal moves and poor strategic play.

One key focus is fraud and criminal activity, whose cost runs into the hundreds of billions of dollars per year or more.  SDS labels many of these criminal activities including money laundering, credit card fraud, check fraud, APP (Authorized Push Payment) fraud (scams), and insurance claims fraud.  As such SDS data provides an attractive foundation for training AI detection models.

Unlike much current activity around synthetic data generation, SDS is not built using large language models.  Instead SDS uses an agent-based virtual world approach.  A key advantage of the SDS design is that all labels are correct:  all fraud is labelled fraud, and only fraud is labelled fraud.  By contrast, much criminal activity is missed in the real world, including 95% of money laundering by a UN estimate.  Hence, even if real data is available, it is often of poor quality for training detection models, or for generating synthetic data.

In practice, access to real data is generally limited to a small number of people at the institution (e.g. a bank) that owns the data.  As such real data provides only a narrow view of activity at a single institution – as opposed to the global view provided by SDS data.  The SDS approach also yields a broad set of synthetic personal information.  This information is highly realistic despite using no information from real individuals.

Development of effective techniques for SDS has required deep expertise across diverse areas.  It has also required significant manual effort.  How to automate some of these efforts remains an open challenge, as do calibration, scaling, and other areas.

Speakers
Erik Altman
IBM Researcher

Watch Event Recording