40 courses found
Comprehensive introduction to predictive modeling, a cornerstone of data science and machine learning. Learn the fundamental concepts, techniques, and tools used to build models while emphasizing both theoretical understanding and practical applications. The topics include we will cover are an in-depth analysis of linear models and different variants, their extension to generalized linear models, and an introduction to nonparametric regression.
This course covers fundamentals of data mining and machine learning within a common statistical framework. Topics include regression, classification, clustering, resampling, regularization, tree-based methods, ensembles, boosting, and Support Vector Machines. Coursework is conducted in the R programming language.
Bayesian inferential methods provide a foundation for machine learning under conditions of uncertainty. Bayesian machine learning techniques can help us to more effectively address the limits to our understanding of world problems. This class covers the major related techniques, including Bayesian inference, conjugate prior probabilities, naive Bayes classifiers, expectation maximization, Markov chain monte carlo, and variational inference. A course covering statistical techniques such as regression.
A graduate-level course on deep learning fundamentals and applications with emphasis on their broad applicability to problems across a range of disciplines. Topics include regularization, optimization, convolutional networks, sequence modeling, generative learning, instance-based learning, and deep reinforcement learning. Students will complete several substantive programming assignments. A course covering statistical techniques such as regression.
Introduces fundamental concepts of computation, data structures, algorithms, & databases, focusing on their role in data science. Covers both theoretical studies & hands-on learning activities. Includes basic data structures, advanced data structures, searching, sorting, greedy algorithms, linear programming, & basics of databases. Will develop computational thinking skills and learn a variety of ways to represent & analyze real-world data.
Covers the fundamentals of probability and stochastic processes. Students will become conversant in the tools of probability, clearly describing and implementing concepts related to random variables, properties of probability, distributions, expectations, moments, transformations, model fit, sampling distributions, discrete and continuous time Markov chains, and Brownian motion.
Introduction to regression modeling. Topics will be discussed first in the context of linear regression, and then revisited in the context of logistic regression, ordinal regression, proportional hazards regression, and random forests. Students will be required to fit the models (both MLE and Bayesian) and use the strategies discussed in class.
Covers data pipeline: techniques to collect data, organize, query & apply the data, and generate products that describe the insights. Topics include Python environments, containers using Docker, data wrangling with pandas, data acquisition via flat files, APIs, JSON formats, and webscraping, relational, document, and graph databases, exploratory data analysis including static & interactive data visualization, dashboards, and cloud computing.
Specialized or advanced topics not in DS current course offerings. Requires (a) approval of the program director and (b) an SDS faculty member who will serve as instructor. Propose a syllabus which includes a week-by-week accounting of the topics, materials (papers and textbooks), and assessments. Reach out to the program director for more details.
Learning tools and concepts for computing on big data. Learn how to use Spark for large-scale analytics and machine learning. Spark is an open-source, general-purpose computing framework that is scalable and blazingly fast. Fundamental data types and concepts will be covered (e.g., resilient distributed datasets, DataFrames) along with Tools for data processing, storage, and retrieval, including Amazon Web Services (AWS).
Covers advanced theoretical concepts for deep neural networks. Topics include convolutional neural networks and their design principles, encoder-decoder architectures, recurrent neural networks, transformers, bounding box detection, image segmentation, generative adversarial networks, diffusion models, etc. Using open-source Python libraries such as NumPy, TensorFlow, and Keras, to understand how theoretical concepts are implemented.
Introduces ways that data and information have historically been constructed in different realms--from medicine to public health to computing--to shed light on the power relationships embedded in some of our present-day and near-future tools, systems, and economic relationships. Will use a historical lens, as well as methods from STS, to give an introduction to how data and power interact in people's lives.
Transition into principal investigators and generators of data science-based knowledge. Develop practical skills necessary to conduct high quality data science research, advance development into producers and critical consumers of research, and further development into professional data scientists broadly defined. Research based career topics covered: time management, research products, types of research positions, and grant writing.
Engages students in identification of a research question, a review of the literature and the application of an existing data science tool or technique (algorithm) to that problem. This is a mentored experience and will allow the student to demonstrate their capacity for research and begin to develop a relationship with a faculty mentor in Data Science. Course requires instructor permission.
PhD level Dissertation Research.