Tensor differentiation
Before going into Machine learning, one of the major concepts is needed tensor calculus. To be more specific we need to know how to differentiate a scaler/vector/matrix with respect to vector/matrix. In this blog I am going discuss that will be needed in to understand Deep learning.
P.S Please don't skip this blog if you don't love mathematics because I will not go deep into mathematics.
$ \newcommand{\norm}[1]{\lVert #1 \rVert} \newcommand{\abs}[1]{| #1 |} \newcommand{\R}{\mathbb{R}} \newcommand{\C}{\mathbb{C}} \newcommand{\ind}{\mathbf{1}} \newcommand{\X}{\mathbf{x}} \newcommand{\A}{\mathbf{a}} \newcommand{\B}{\mathbf{b}} \newcommand{\Z}{\mathbf{z}} \newcommand{\ip}[2]{\langle #1, #2 \rangle} \newcommand{\NL}{\linebreak{NL}} \newcommand{\*}{\times} \newcommand{\pd}{\partial} $Differentiation of scalar with respect to vector/Matrix
Let the function is $f : \R^n \longrightarrow \R$,
The function is taking a input of n scalar variable(real value) and output a single scalar real value.
let the n variable of the function is $x_1,x_2,... x_n$.
So, $\mathbf {x}=\begin{bmatrix} x_1\\x_2\\x_3\\ ..\\.. \\ x_n\end{bmatrix} $ ; $ {\pd f(\X) \over \pd \X } = \begin{bmatrix} \pd f(\X) \over \pd x_1 \\ \pd f(\X) \over \pd x_2 \\ ..\\ .. \\ \pd f(\X) \over \pd x_n \end{bmatrix} $
$\X$ is a vector of length n. We can represent it as $n\* 1$ matrix. Derivative of $f(\X)$ with respect to $\X$, i.e $\pd f(\X) \over\pd \X$ will be a vector of size $n\*1$ when you do derivative of a scalar with respect to vector the output is always a vector as the size of vector by which you are doing derivative.$Note:$ Similarly when you do derivative of a scalar with respect to matrix the output will always be a matrix as the size of the matrix by which you are doing derivative. You can visualize all this if you have done course in linear algebra Matrix calculus.
First let us see the Matrix algebra which will help us to workout in practical
Let $ \A,\X,\B $ is a vector of size in the form $n\* 1$ . $A,X,C$ is matrix of size in the form $n\* m$.
Note that"In the form" doesn't mean exactly, like in $\A^TX\B$, $\A$ may be of size $n_1\* 1$ and $\B$ of size $n_2\* 1$ , as $\A^TX\B$ is scalar and matrix multiplication property must hold so size of $X$ is $n_1\* n_2$ .
So $\X^Ta ,  \A^T\X,  \X^T\A\X,$   $ \A^TX\B,  \A^TX^T\B ,  \A^TX^TCX\B,  (X\B+\B)^TC(X\A+\B)$ are all scalar.
Scalar with respect to Vector
- ${\pd \A^T\X \over \pd \X} =\A$ as $\A^T\X=\X^T\A$=scalar so ${\pd \X^T\A \over \pd \X} =\A$
- ${\pd \X^T A \X \over \pd X} = (A+A^T) \X $ Note: $\X^TA\X$ is quadratic form of equation. By using this two we can do derivative of any scalar in matrix form
Scalar with respect to Matrix
- ${\pd \A^TX\B \over \pd X} = \A\B^T $
- ${\pd (\A^TX^T\B) \over \pd X} = {\pd (\A^TX^T\B)^T \over \pd X} ={\pd ((X^T\B)^T\A) \over \pd X} = {\pd (\B^TX\A) \over \pd X}= \B\A^T $ as $\A^TX^T\B$ is scalar so $\A^TX^T\B=(\A^TX^T\B)^T $
- ${\pd \A^TX^TCX\B \over \pd X} = C^TX \A \B^T +CX \B a^T$
- ${\pd (X\A+\B)^TC(X\A+\B) \over \pd X} = {\pd ((X\A)^T+\B^T)C(X\A+\B) \over \pd X}={\pd (\A^TX^T+\B^T)(CX\A+C\B)) \over \pd X}$ $={{\pd \A^TX^TCX\A \over \pd X} + {\pd \A^TX^TC\B \over \pd X} + {\pd \B^TCX\A \over \pd X} +{\pd \B^TC\B \over \pd X }}= {C^TX \A \A^T +CX \A \A^T+C\B\A^T+ (\B^TC)^T \A^T +0 } $ $ = { (C^TX\A + C\B +CX \A +C^T\B)\A^T} = {(C+ C^T)(X\A+\B)\A^T } $
Vector/matrix with respect to vector
When you derivative a vector with respect to a vector out come will be a matrix. For example if $\X$ has shape of $n\* 1$ and $\Z$ has shape of $m\* 1$ then $ {\pd \X^T \over \pd \Z}$ is a matrix of shape $m\*n$.
Trick to remember(No mathematical logic): Derivative outcome shape is like multiplying to matrix.for the above example $\Z\X^T$ i.e $m\* 1$ to $1\* n$ so the output shape is $m\*n$.
Vector with respect to vector
- ${\pd \X^T \over \pd \X}= I $
- ${\pd \X^TA \over \pd \X}= A $
- $ {\pd (A\X) \over \pd \Z}= { A {\pd \X \over \pd \Z}} $
Matrix with respect to vector
- ${ \pd X^{-1}\over \pd \Z } ={ X^{-1}{\pd X\over \pd \Z }X^{-1}} $
- ${\pd AXB \over \pd \Z}= A{\pd X \over \pd \Z}B $
- ${\pd XY \over \pd \Z}= {X{\pd Y \over \pd \Z} + {\pd X \over \pd \Z}Y} $
Further reading
Now you will be able to understand the problems which are formulated with vector/matrix and optimized. For more reading about these topics in more detail(not in shortcut ;) ) here I have given some resource.