Search results for Pattern Recognition and Machine Learning

9 - Nonsmooth Optimization Methods
Stephen J. Wright, University of Wisconsin, Madison, Benjamin Recht, University of California, Berkeley
Book:

Optimization for Data Analysis

Published online:

31 March 2022

Print publication:

21 April 2022, pp 153-169
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation
Summary

Here, we describe algorithms for minimizing nonsmooth functions and composite nonsmooth functions, which are the sum of a smooth function and a (usually elementary) nonsmooth function. We start with the subgradient descent method, whose search direction is the minimum-norm element of the subgradient. We then discuss the subgradient method, which steps along an arbitrary direction drawn from the subdifferential. Next, we describe proximal-gradient algorithms for nonsmooth composite optimization, which make use of the gradient of the smooth part of the function and the proximal operator associated with the nonsmooth part. Finally, we describe the proximal point method, a framework optimization that is valuable both as a fundamental method in its own right and as a building block for the augmented Lagrangian approach described in the next chapter.

11 - Differentiation and Adjoints
Stephen J. Wright, University of Wisconsin, Madison, Benjamin Recht, University of California, Berkeley
Book:

Optimization for Data Analysis

Published online:

31 March 2022

Print publication:

21 April 2022, pp 188-199
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation
Summary

First derivatives (gradients) are needed for most of the algorithms described in the book. Here, we describe how these gradients can be computed efficiently for functions that have the form of arising in deep learning. The reverse mode of automatic differentiation, often called “back-propagation” in the machine learning community, is described for several problems with nested-composite and progressive structure that arises in neural network training. We provide another perspective on these techniques, based on a constrained optimization formulation and optimality conditions for this formulation.

8 - Nonsmooth Functions and Subgradients
Stephen J. Wright, University of Wisconsin, Madison, Benjamin Recht, University of California, Berkeley
Book:

Optimization for Data Analysis

Published online:

31 March 2022

Print publication:

21 April 2022, pp 132-152
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation
Summary

Here, we define subgradients and subdifferentials of nonsmooth functions. These are a generalization of the concept of gradients for smooth functions, that can be used as the basis of algorithms. We relate subgradients to directional derivatives and to the normal cones associated with convex sets. We introduce composite nonsmooth functions that arise in regularized optimization formulations of data analysis problems and describe optimality conditions for minimizers of these functions. Finally, we describe proximal operators and the Moreau envelope, objects associated with nonsmooth functions that are the basis of algorithms for nonsmooth optimization described in the next chapter.

3 - Descent Methods
Stephen J. Wright, University of Wisconsin, Madison, Benjamin Recht, University of California, Berkeley
Book:

Optimization for Data Analysis

Published online:

31 March 2022

Print publication:

21 April 2022, pp 26-54
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation
Summary

In this section, we discuss fundamental methods, mostly based on gradient information, that yield descent, that is, the function value decreases at each iteration. We start with the most basic method, the steepest-descent method, analyzing its convergence under different convexity/nonconvexity assumptions on the objective function. We then discuss more general descent methods, based on descent directions other than the negative gradient, showing conditions on the search direction and the steplength that allow convergence results to be proved. We also discuss a method that also makes use of Hessian information, showing that it can find a point satisfying approximate second-order optimality conditions and finding an upper bound on the number of iterations required to do so. We then discuss mirror descent, a class of gradient methods based on more general distance metrics that are particularly useful in optimizing over the unit simplex – a problem that arises often in data science. We conclude by discussing the PL condition, a generalization of the strong convexity condition that allows linear convergence rates to be proved.

Frontmatter
Stephen J. Wright, University of Wisconsin, Madison, Benjamin Recht, University of California, Berkeley
Book:

Optimization for Data Analysis

Published online:

31 March 2022

Print publication:

21 April 2022, pp i-iv
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

Optimization for Data Analysis

Stephen J. Wright, Benjamin Recht
Published online:

31 March 2022

Print publication:

21 April 2022
- Book
- - Get access
    
    Buy a print copy
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation
Optimization techniques are at the core of data science, including data analysis and machine learning. An understanding of basic optimization techniques and their fundamental properties provides important grounding for students, researchers, and practitioners in these areas. This text covers the fundamentals of optimization algorithms in a compact, self-contained way, focusing on the techniques most relevant to data science. An introductory chapter demonstrates that many standard problems in data science can be formulated as optimization problems. Next, many fundamental methods in optimization are described and analyzed, including: gradient and accelerated gradient methods for unconstrained optimization of smooth (especially convex) functions; the stochastic gradient method, a workhorse algorithm in machine learning; the coordinate descent approach; several key algorithms for constrained optimization problems; algorithms for minimizing nonsmooth functions arising in data science; foundations of the analysis of nonsmooth functions and optimization duality; and the back-propagation approach, relevant to neural networks.

Bibliography
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp 327-334
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

3 - Basic Parametric Models and a Statistical Perspective on Learning
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp 37-62
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

1 - Introduction
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp 1-12
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

6 - Neural Networks and Deep Learning
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp 133-162
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

Acknowledgements
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp ix-x
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

2 - Supervised Learning: A First Approach
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp 13-36
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

Frontmatter
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp i-iv
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

11 - User Aspects of Machine Learning
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp 287-308
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

Notation
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp xi-xii
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

Contents
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp v-viii
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

4 - Understanding, Evaluating, and Improving Performance
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp 63-90
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

5 - Learning Parametric Models
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp 91-132
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

8 - Non-linear Input Transformations and Kernels
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp 189-216
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

10 - Generative Models and Learning from Unlabelled Data
Andreas Lindholm, Niklas Wahlström, Uppsala Universitet, Sweden, Fredrik Lindsten, Linköpings Universitet, Sweden, Thomas B. Schön, Uppsala Universitet, Sweden
Book:

Machine Learning

Published online:

27 May 2022

Print publication:

31 March 2022, pp 247-286
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation

Pattern Recognition and Machine Learning

Refine search

Refine search

Actions for selected content:

2190 results in Pattern Recognition and Machine Learning

9 - Nonsmooth Optimization Methods

Summary

11 - Differentiation and Adjoints

Summary

8 - Nonsmooth Functions and Subgradients

Summary

3 - Descent Methods

Summary

Frontmatter

Optimization for Data Analysis

Bibliography

3 - Basic Parametric Models and a Statistical Perspective on Learning

1 - Introduction

6 - Neural Networks and Deep Learning

Acknowledgements

2 - Supervised Learning: A First Approach

Frontmatter

11 - User Aspects of Machine Learning

Notation

Contents

4 - Understanding, Evaluating, and Improving Performance

5 - Learning Parametric Models

8 - Non-linear Input Transformations and Kernels

10 - Generative Models and Learning from Unlabelled Data

Pattern Recognition and Machine Learning

Refine search

Refine search

Actions for selected content:

Save Search

2190 results in Pattern Recognition and Machine Learning

Summary

Summary

Summary

Summary

Optimization for Data Analysis