Skip to content
markovka17Public

Latest commit

 

History

201 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

logo5v1

Deep Learning for Audio (DLA)

  • Lecture and seminar materials for each week are in ./week* folders, see README.md for materials and instructions
  • Any technical issues, ideas, bugs in course materials, contribution ideas - add an issue
  • The current version of the course is conducted in autumn 2026 at the CS Faculty of HSE.

For previous years versions, see Past Versions section.

Syllabus

  • week01 Introduction to Course

    • Lecture: Introduction to Course + Inspiration
    • Seminar: Free talk
    • Self-Study: Introduction to PyTorch and basic devOps
  • week02 Introduction to Digital Signal Processing

    • Lecture: Signals, Fourier Transform, spectrograms, MelScale, MFCC
    • Seminar: DSP in practice, spectrogram creation, IRF, frequency filtering
  • week03 Automatic Speech Recognition I

    • Lecture: Metrics, Datasets, Connectionist Temporal Classification (CTC), DeepSpeech2, Conformer, Beam Search, Language models
    • Seminar: Audio Augmentations, WER and CER, CTC Decoding
  • week04 Automatic Speech Recognition II

    • Lecture: LAS, Hybrid CTC/Attention, OpenAI Whisper, RNN-T, Streaming ASR, Decoder-only ASR
    • Seminar: Whisper: greedy decoding, prompting, alignment in cross-attention, language forcing

TBA

Homeworks and Projects

TBA

See our project template.

Resources

Some of the weeks have English recordings. See the corresponding sub-directories.

Contributors & course staff

Course materials and teaching (in different years) were delivered by:

Past Versions

Releases

Packages

Contributors

Languages