Supported by CCR Office of Science and Technology Resources (OSTR)
ncibtep@nih.gov

Bioinformatics Training and Education Program

Upcoming Classes & Events

August

Organized by
Center of Excellence in Immunology
Description

This two-day national symposium addresses recent advances in the field and should be an exciting forum for discussion and debate on the current understanding of cancer immunology in the era of omics and artificial intelligence.

Confirmed Speakers:

  • Grégoire Altan-Bonnet, NCI
  • Avinash Bhandoola, NCI
  • Remy Bosselut, NCI
  • Mary Carrington, NCI
  • Leah Cook, NCI
  • Amiran Dzutsev, NCIRead More

This two-day national symposium addresses recent advances in the field and should be an exciting forum for discussion and debate on the current understanding of cancer immunology in the era of omics and artificial intelligence.

Confirmed Speakers:

  • Grégoire Altan-Bonnet, NCI
  • Avinash Bhandoola, NCI
  • Remy Bosselut, NCI
  • Mary Carrington, NCI
  • Leah Cook, NCI
  • Amiran Dzutsev, NCI
  • Donna Farber, Columbia University
  • Paul François, Université de Montréal
  • Romina Goldszmid, NCI
  • Timothy Greten, NCI
  • Peng Jiang, NCI
  • Yann LeCun, New York University
  • Lichun Ma, NCI
  • Bali Pulendran, Stanford School of Medicine
  • Barbara Reherman, NIDDK
  • Eytan Ruppin, Cedars-Sinai Medical Center
  • Eldad Shulman, Cedars-Sinai Medical Center
  • Naomi Taylor, NCI
  • Giorgio Trinchieri, NCI
  • John Tsang, Yale University
  • Roxane Tussiwand, NCI 
  • Golnaz Vahedi, University of Pennsylvania School of Medicine
  • Roberto Weigert, NCI
  • Ramnik Xavier, Harvard University
  • Li Yang, NCI
  • Chen Zhao, NCI
  • Marlies Meisel, University of Pittsburgh School of Medicine
  • Rosandra Kaplan, NCI

Main Topics

  • DATA SCIENCE AND DEEP LEARNING IN CANCER IMMUNITY
  • TUMOR MICROENVIRONMENT
  • MICROBIOME AND CANCER
  • T CELLS IN CANCER IMMUNITY
Organized by
CIT Technology Training Program
Description
You know the basics of prompt engineering—but great AI results require more than writing better prompts. In this course, you'll learn the advanced techniques that separate casual AI users from AI power users. Discover how to refine and troubleshoot prompts, guide AI through complex tasks, structure outputs for higher quality, and use AI as a strategic thinking partner rather than just a content generator. Through practical NIH-focused examples and hands-on exercises, you'll explore Read More
You know the basics of prompt engineering—but great AI results require more than writing better prompts. In this course, you'll learn the advanced techniques that separate casual AI users from AI power users. Discover how to refine and troubleshoot prompts, guide AI through complex tasks, structure outputs for higher quality, and use AI as a strategic thinking partner rather than just a content generator. Through practical NIH-focused examples and hands-on exercises, you'll explore prompt optimization, multi-step prompting, critical thinking frameworks, and methods for improving accuracy, clarity, and usefulness.
Join Meeting
Organized by
BTEP
Description

Presenting the third and final event in a 3-part series on Project and Data Management; created and presented by the NIDDK Biostatistics Program and the Office of the Clinical Director.

Learning Objectives:  

  1. To understand best practices for research safety monitoring and reporting
  2. To reflect on the importance of ongoing data review to ensure data integrity
  3. Read More

Presenting the third and final event in a 3-part series on Project and Data Management; created and presented by the NIDDK Biostatistics Program and the Office of the Clinical Director.

Learning Objectives:  

  • To understand best practices for research safety monitoring and reporting
  • To reflect on the importance of ongoing data review to ensure data integrity
  • To review successful strategies for timely reporting of clinical data in CT.gov for regulatory compliance 
  •  

    September

    Join Meeting
    Organized by
    Leidos Biomedical Research (LBR) Frederick National Lab for Cancer Research (FNLCR)
    Description

    No one fights cancer alone — not patients, and not scientists. Every discovery at the Frederick National Laboratory for Cancer Research builds on the shared efforts of researchers, clinicians, and computational biologists working towards a common goal. From high-throughput sequencing of complex experimental designs to downstream computational analysis, this talk highlights how the Center for Cancer Research’s Collaborative Bioinformatics Resource (CCBR) team leverages high-performance computing to transform raw sequencing data into meaningful biological insights. Read More

    No one fights cancer alone — not patients, and not scientists. Every discovery at the Frederick National Laboratory for Cancer Research builds on the shared efforts of researchers, clinicians, and computational biologists working towards a common goal. From high-throughput sequencing of complex experimental designs to downstream computational analysis, this talk highlights how the Center for Cancer Research’s Collaborative Bioinformatics Resource (CCBR) team leverages high-performance computing to transform raw sequencing data into meaningful biological insights. This behind-the-scenes view of collaborative,data-driven genomics highlights how rigorous analysis and reproducible accelerates discoveries that move us closer to better prevention, diagnosis, and treatment.

    Join Meeting
    Organized by
    NIH Library
    Description

    This one-hour online training provides researchers with an overview of online resources for locating research datasets, data repositories, and data publications for data sharing and re-use. Participants will learn search strategies for locating datasets through federated data search portals and generalist data repositories, including directories for locating discipline-specific and institutional data repositories. An overview of key issues to consider when re-using datasets or when locating a data repository for sharing Read More

    This one-hour online training provides researchers with an overview of online resources for locating research datasets, data repositories, and data publications for data sharing and re-use. Participants will learn search strategies for locating datasets through federated data search portals and generalist data repositories, including directories for locating discipline-specific and institutional data repositories. An overview of key issues to consider when re-using datasets or when locating a data repository for sharing and preservation purposes will be discussed. 

    By the end of this training, attendees will be able to:  

    • Locate different types of data repositories and datasets 

    • Identify issues to consider with data repositories 

    • Discuss how data repositories can improve reproducibility
    • Identify issues to consider when re-using datasets 

    • Describe guidelines and resources for citing datasets 

    Attendees are not expected to have any prior knowledge of these resources to be successful in this training. 

    Organized by
    NIH Library
    Description

    Claude 101 is part 1 of a two-part series. 

    This hour and half online training led by Anthropic will cover the fundamentals of using Claude effectively in your daily NIH workflows. Attendees will learn to navigate the Claude interface, apply best practices for prompt writing, and utilize key features such as working with documents, Projects, and Artifacts. The training will also demonstrate real-world use cases Read More

    Claude 101 is part 1 of a two-part series. 

    This hour and half online training led by Anthropic will cover the fundamentals of using Claude effectively in your daily NIH workflows. Attendees will learn to navigate the Claude interface, apply best practices for prompt writing, and utilize key features such as working with documents, Projects, and Artifacts. The training will also demonstrate real-world use cases relevant to NIH staff for improving productivity, and highlight security and responsible-use considerations tailored for federal environments. 

     By the end of this training, attendees will be able to: 

    • Navigate the Claude interface and use foundational features, including working with documents, Projects, and Artifacts. 

    • Apply effective prompting strategies to generate accurate, useful outputs for NIH-specific tasks. 

    • Identify everyday NIH use cases and understand best practices for responsible use of generative AI tools like Claude. 

    Attendees are not expected to have any prior knowledge of the tool to be successful in this training. 

    Organized by
    NIH Library
    Description

    Claude 201 is part 2 of a two-part series. 

    This hour and half online training led by Anthropic will dive deeper into intermediate and advanced strategies for maximizing Claude in NIH workflows. Building on the fundamentals from Claude 101, this training will focus on structured and multi-step prompting, working effectively with longer documents and datasets, and Read More

    Claude 201 is part 2 of a two-part series. 

    This hour and half online training led by Anthropic will dive deeper into intermediate and advanced strategies for maximizing Claude in NIH workflows. Building on the fundamentals from Claude 101, this training will focus on structured and multi-step prompting, working effectively with longer documents and datasets, and using Projects to organize ongoing work and build reusable context. Attendees will also learn how to integrate Claude into specialized NIH tasks and optimize outputs for research, administrative, and policy workflows. 

    By the end of this training, attendees will be able to: 

    • Use structured and multi-step prompting techniques to handle complex tasks and improve output quality. 

    • Work effectively with documents, longer-form content, and data inside Claude to support research and analysis workflows. 

    • Set up and use Projects to organize ongoing work, build reusable context, and collaborate on NIH-specific initiatives. 

    Attendees are expected to be familiar with the basic functions of Claude to be successful in this training (gained by attending Claude 101, attending another relevant training, and/or using Claude previously). 

    Join Meeting
    Organized by
    BTEP
    Description

    Partek Flow is a point-and-click platform for building analysis workflows for Next Generation Sequences (NGS), including DNA, bulk and single-cell RNA, spatial transcriptomics, ATAC, and ChIP, helping scientists avoid the steep learning curve of code-based NGS analysis. This class is demonstration-only. Starting from single cell RNA expression matrix, Illumina scientist will illustrate how to conduct QC, perform cell type classification, obtain differential expression results, and generate visualizations. No prior experience or access to Partek Read More

    Partek Flow is a point-and-click platform for building analysis workflows for Next Generation Sequences (NGS), including DNA, bulk and single-cell RNA, spatial transcriptomics, ATAC, and ChIP, helping scientists avoid the steep learning curve of code-based NGS analysis. This class is demonstration-only. Starting from single cell RNA expression matrix, Illumina scientist will illustrate how to conduct QC, perform cell type classification, obtain differential expression results, and generate visualizations. No prior experience or access to Partek Flow is required. Attendance is limited to NIH staff.

    Organized by
    NIH Library
    Description

    This one-hour online training, is the first of a two-part series, which introduces participants to cleaning and exploring a patient health dataset using Python and pandas. Attendees will load tabular data, inspect structure and data types, summarize columns, and identify common data quality problems such as missing values, inconsistent formats, and duplicate records. They will then apply practical fixes, including standardizing height and weight units, parsing and normalizing dates of birth, splitting combined fields, Read More

    This one-hour online training, is the first of a two-part series, which introduces participants to cleaning and exploring a patient health dataset using Python and pandas. Attendees will load tabular data, inspect structure and data types, summarize columns, and identify common data quality problems such as missing values, inconsistent formats, and duplicate records. They will then apply practical fixes, including standardizing height and weight units, parsing and normalizing dates of birth, splitting combined fields, and using Boolean masks to flag or correct implausible values.​

    By the end of this session students will be able to:

    • Import CSV data into pandas DataFrames and quickly understand column types, basic statistics, and overall data quality.​
    • Identify duplicate or repeated patient records and decide whether to keep, correct, or remove them.​
    • Detect and handle missing or inconsistent values using methods such as isna, fillna, filtering, and conditional replacement.​
    • Standardize mixed formats (for example, heights with and without units, date strings in different formats, and numeric values embedded in text).​
    • Create derived columns such as systolic and diastolic blood pressure, and use logical conditions to flag questionable or out-of-range values.​

    Attendees are expected to have:

    • Basic Python coding knowledge
    • Familiarity with an IDE and loading script and data files into the IDE. (Colab, Jupyter Notebooks) 

    Requirements: 

    • Participants will receive a script file and data files prior to the training. These should be loaded and ready to use before the training session begins.
    Organized by
    NIH Library
    Description

    This one-hour online training shows attendees how to use generative AI to accelerate scientific discovery and streamline data analysis. This training, open to all disciplines, demonstrates how AI-assisted coding can quickly turn ideas into functional analysis tools with minimal manual effort.  

    By the end of this training, attendees will be able to:  
    • Understand how MATLAB supports low-code, reproducible research while Read More

    This one-hour online training shows attendees how to use generative AI to accelerate scientific discovery and streamline data analysis. This training, open to all disciplines, demonstrates how AI-assisted coding can quickly turn ideas into functional analysis tools with minimal manual effort.  

    By the end of this training, attendees will be able to:  
    • Understand how MATLAB supports low-code, reproducible research while remaining flexible for advanced customization using MATLAB Copilot  
    • Apply generative AI tools to accelerate data analysis while maintaining scientific rigor and reproducibility  
    • Integrate generative AI into existing MATLAB workflows to reduce development time while preserving transparency and control  

    Attendees should be familiar with basic MATLAB functions to succeed in this training.

    Organized by
    NIH Library
    Description

    This one-hour online training, the second session of the two-part series,  focuses on reshaping and enriching the cleaned patient dataset to prepare it for analysis and reporting. Attendees will practice splitting and recombining columns (for example, separating full names into first and last names), converting columns to appropriate data types, and engineering new fields such as outlier indicators and blood pressure status labels. The session also covers merging multiple tables (patient details, contact Read More

    This one-hour online training, the second session of the two-part series,  focuses on reshaping and enriching the cleaned patient dataset to prepare it for analysis and reporting. Attendees will practice splitting and recombining columns (for example, separating full names into first and last names), converting columns to appropriate data types, and engineering new fields such as outlier indicators and blood pressure status labels. The session also covers merging multiple tables (patient details, contact information, and subsets of records) and filtering or subsetting data to answer specific analytical questions.​

    By the end of this training, attendees will be able to:

    • Reshape and restructure data by splitting and combining columns, changing data types, and reordering or selecting relevant fields.​
    • Engineer clinically useful features, including z-score–based outlier flags, hypertension indicators, and combined status columns for downstream models or dashboards.​
    • Merge and join DataFrames using common keys (such as patient ID) to bring together core data with supplemental tables like contact information.​
    • Filter and subset records based on multiple conditions (for example, patients with diabetes and abnormal blood pressure) to create analysis-ready datasets.​

    Attendees are expected to have:

    • To have attended Intro to Data Wrangling Using Python - Part 1 of the series
    • Basic Python coding knowledge

    Familiarity with an IDE and loading script and data files into the IDE. (Colab, Jupyter Notebooks) 

    Requirements: 

    • Participants will receive a script file and data files prior to the training. These should be loaded and ready to use before the training session begins. 
    Organized by
    NIH Library
    Description

    This one-hour online training introduces attendees to modeling and simulation of biological systems using MATLAB’s SimBiology and BioPipeline Designer toolboxes. SimBiology is a versatile toolbox for modeling, simulating, and analyzing dynamic biological systems such as metabolic pathways, signaling cascades, and pharmacokinetics/pharmacodynamics (PK/PD) models. BioPipeline Designer complements this by streamlining workflows for integrating biological data and automating computational analyses. 

    By Read More

    This one-hour online training introduces attendees to modeling and simulation of biological systems using MATLAB’s SimBiology and BioPipeline Designer toolboxes. SimBiology is a versatile toolbox for modeling, simulating, and analyzing dynamic biological systems such as metabolic pathways, signaling cascades, and pharmacokinetics/pharmacodynamics (PK/PD) models. BioPipeline Designer complements this by streamlining workflows for integrating biological data and automating computational analyses. 

    By the end of this training, attendees will be able to: 

    • Describe the capabilities and applications of SimBiology and BioPipeline Designer for modeling and analyzing biological systems. 

    • Construct and parameterize basic models of biological processes using SimBiology’s graphical and programmatic interfaces. 

    • Simulate dynamic behaviors of biological systems, such as time-course analyses, and interpret simulation results. 

    • Automate and streamline data integration workflows using BioPipeline Designer to enhance reproducibility and efficiency. 

    • Access and utilize resources for further learning, including tutorials, user guides, and MATLAB community forums 

    Attendees are expected to be familiar with the basic functions of the MATLAB to be successful in this training. 

    Organized by
    NIH Library
    Description

    This 90-minute online roundtable explores practical applications of artificial intelligence (AI) in statistics and data analysis across the NIH research landscape. Brief presentations from panelists representing statistical, data science, and research perspectives will be followed by an open moderated discussion. Attendees will come away able to identify real-world AI use cases in research workflows, describe the opportunities and limitations of AI-assisted methods, and discuss how AI may shape the future of statistical practice, biomedical Read More

    This 90-minute online roundtable explores practical applications of artificial intelligence (AI) in statistics and data analysis across the NIH research landscape. Brief presentations from panelists representing statistical, data science, and research perspectives will be followed by an open moderated discussion. Attendees will come away able to identify real-world AI use cases in research workflows, describe the opportunities and limitations of AI-assisted methods, and discuss how AI may shape the future of statistical practice, biomedical research, and decision-making at NIH and HHS. The discussion will also touch on considerations of bias, reproducibility, and responsible AI use within federally-funded research contexts.

    October

    Organized by
    NIH Library
    Description

    This one-hour online training will cover the fundamentals, applications, and ethical considerations of Artificial Intelligence (AI). Attendees will explore key topics such as machine learning, deep learning, data handling, and real-world AI applications across various industries. The session will also delve into the ethical implications of AI and provide insights on becoming AI literate. Whether you're a seasoned professional or just starting your AI journey, this session will equip you with essential knowledge to Read More

    This one-hour online training will cover the fundamentals, applications, and ethical considerations of Artificial Intelligence (AI). Attendees will explore key topics such as machine learning, deep learning, data handling, and real-world AI applications across various industries. The session will also delve into the ethical implications of AI and provide insights on becoming AI literate. Whether you're a seasoned professional or just starting your AI journey, this session will equip you with essential knowledge to navigate the AI landscape effectively and make informed decisions in our data-driven world.

    By the end of this training, attendees will be able to: 

    • Understand the core concepts of AI 
    • Recognize the significance of ethical considerations in AI 
    • Begin the journey toward AI literacy

    Attendees are not expected to have any prior knowledge of AI to be successful in this training.

    Organized by
    NIH Library
    Description

    This hour and a half online training covers how to analyze and model data using interactive tools in MATLAB. Through live demonstrations and examples, attendees will learn to solve many steps in a data analysis workflow without writing any code. The interactive tools can generate the MATLAB code needed to reproduce the work programmatically. 

    By the end of this training, attendees will be able to:

    • Use interactive tools Read More

    This hour and a half online training covers how to analyze and model data using interactive tools in MATLAB. Through live demonstrations and examples, attendees will learn to solve many steps in a data analysis workflow without writing any code. The interactive tools can generate the MATLAB code needed to reproduce the work programmatically. 

    By the end of this training, attendees will be able to:

    • Use interactive tools for data visualization, cleaning, and modeling
    • Automatically generate code to replicate interactive work
    • Capture work in easy-to-write scripts and functions
    • Share results by automatically creating reports

    This training taught by MathWorks. Attendees are not expected to have any prior knowledge of MATLAB, but experienced users will also benefit from new tools, tips, and tricks from the latest releases. This training is an introductory level; no software installation required.

    Organized by
    NIH Library
    Description

    This one hour and half hour online training will equip attendees with essential knowledge and skills for effective interactions with Large Language Model (LLM) AI chatbots. Explore the intricacies of prompt engineering and its pivotal role in optimizing the conversational capabilities of LLMs. Emphasizing best practices and practical applications, this training features live demonstrations and provides valuable skills for the effective use of LLMs. 

    This one hour and half hour online training will equip attendees with essential knowledge and skills for effective interactions with Large Language Model (LLM) AI chatbots. Explore the intricacies of prompt engineering and its pivotal role in optimizing the conversational capabilities of LLMs. Emphasizing best practices and practical applications, this training features live demonstrations and provides valuable skills for the effective use of LLMs. 

    By the end of this training, attendees will be able to:  

    • Define LLMs, prompt patterns, and prompt engineering
    • Identify potential uses and issues to consider when using LLMs in the biomedical research field
    • Use a selection of prompt patterns to improve generated output from LLMs
    • Identify resources for learning more about prompt engineering in LLMs 

    Attendees are not expected to have any prior knowledge of AI chatbots to be successful in this training. 

    Join Meeting
    Organized by
    BTEP
    Description

    Partek Flow is a point-and-click platform for building analysis workflows for Next Generation Sequences (NGS), including DNA, bulk and single-cell RNA, spatial transcriptomics, ATAC, and ChIP, helping scientists avoid the steep learning curve of code-based NGS analysis. In this demonstration-only class, an Illumina scientist will show a bulk ATAC-sequencing workflow starting from FASTQ files through peak and motif detection as well as comparison of peaks found across samples. No prior experience or access to Read More

    Partek Flow is a point-and-click platform for building analysis workflows for Next Generation Sequences (NGS), including DNA, bulk and single-cell RNA, spatial transcriptomics, ATAC, and ChIP, helping scientists avoid the steep learning curve of code-based NGS analysis. In this demonstration-only class, an Illumina scientist will show a bulk ATAC-sequencing workflow starting from FASTQ files through peak and motif detection as well as comparison of peaks found across samples. No prior experience or access to Partek Flow is required. Attendance is limited to NIH staff.

    Organized by
    NIH Library
    Description

    This 45-minute online Lunch and Learn training will help attendees develop their own customized strategy for responsibly incorporating generative artificial intelligence (AI) tools, such as ChatGPT, into their workflows. 

    By the end of this training, attendees will be able to: 

    • Assess appropriate use cases for generative AI tools within their specific research/work context&Read More

    This 45-minute online Lunch and Learn training will help attendees develop their own customized strategy for responsibly incorporating generative artificial intelligence (AI) tools, such as ChatGPT, into their workflows. 

    By the end of this training, attendees will be able to: 

    • Assess appropriate use cases for generative AI tools within their specific research/work context 

    • Develop a customized generative AI usage strategy 

    • Document their approach for using generative AI tools 

    Attendees are not expected to have any prior knowledge of generative AI tools to be successful in this training. 

    Description

    This 30-minute online training provides a high-level overview of recent developments in artificial intelligence (AI). Each session highlights emerging trends, tools, and use cases in the evolving AI landscape, with an emphasis on practical relevance and responsible use. Whether you're just getting started or looking to stay current, this training offers timely insights in a concise format.  

    By the end of this Read More

    This 30-minute online training provides a high-level overview of recent developments in artificial intelligence (AI). Each session highlights emerging trends, tools, and use cases in the evolving AI landscape, with an emphasis on practical relevance and responsible use. Whether you're just getting started or looking to stay current, this training offers timely insights in a concise format.  

    By the end of this training, attendees will be able to:   

    • Summarize key trends and developments in AI 

    • Identify new tools, capabilities, or applications relevant to their work 

    • Describe considerations for ethical and responsible use of AI technologies 

    Attendees are not expected to have any prior knowledge to be successful in this training.