next previous
Up: A statistical method for


3 Some properties

Any statistical test depends on the data made of random variables. Therefore, it is necessary to know the variance of these random variables.

Throughout this paper, we only consider the cases for which random variables vary continuously. In this kind of cases, the cumulative distribution function can be expressed as
\begin{displaymath}
f(x)=\int_{-\infty }^x\rho (x){\rm d}x, \end{displaymath} (10)
where $\rho (x)$ is the density function of the distribution (or the distribution density). $\rho (x)$ has the following properties:

1 $\rho (x)\geq 0$;

2 $\int_{-\infty }^\infty \rho (x){\rm d}x=1$;

3 $p\{a\leq \xi \leq b\}=f(b)-f(a)=\int_a^b\rho (x){\rm d}x$.

According to (10) and property 1, we have
\begin{displaymath}
\frac{{\rm d}f}{{\rm d}x}=\rho (x)\geq 0. \end{displaymath} (11)
We can limit our discussion to the cases where $\rho (x)\neq 0$. The reason is that, when $\rho (x)=0$, $\{\xi =x\}$ is an impossible event. A sample containing such events will not be adopted, or for such a sample, the distribution will not be applied. Therefore, in the concerned intervals, f(x) must be a continuous and monotonic function of x. Thus, in these intervals, there must exist a continuous and monotonic inverse function


x=x(f)

(12)

and its derivative
\begin{displaymath}
\frac{{\rm d}x}{{\rm d}f}=\frac 1{\frac{{\rm d}f}{{\rm d}x}}=\frac 1{\rho (x)}. \end{displaymath} (13)
According to the error propagation law, we have
\begin{displaymath}
{\rm Var}\{x\}=\left(\frac{{\rm d}x}{{\rm d}f}\right)^2{\rm Var}\{f\}=\frac 1{\rho ^2(x)}{\rm Var}\{f\} \end{displaymath} (14)
in these intervals.

The deviation of the distribution function $f_N\left( x\right) $of a sample from the given cumulative distribution function f(x), $\left\vert
f_N(x)-f(x)\right\vert $, can be considered to be caused by the deviation of the random variable, xi, $i\in I$, from its expected value, x0i, $i\in I$, where $I\equiv \{i\vert 1\leq i\leq N\}$. Let
\begin{displaymath}
x_{0i}=x[f_N(x_i)],\qquad \qquad \qquad \forall i\in I. \end{displaymath} (15)
Then
\begin{displaymath}
f(x_{0i})=f_N(x_i),\qquad \qquad \qquad \forall i\in I. \end{displaymath} (16)
According to the definition of the distribution function, we have
\begin{displaymath}
f_N(x_{0i})=f_N(x_i),\qquad \qquad \qquad \forall i\in I. \end{displaymath} (17)
Therefore
\begin{displaymath}
f_N(x_{0i})=f(x_{0i}),\qquad \qquad \qquad \forall i\in I. \end{displaymath} (18)
This shows that the value of the distribution function of sample $S_0\equiv \{x_{0i}\vert i\in I\}$ at any data point x0i, fN(x0i), meets exactly its expected value, f(x0i), $i\in I$. The distribution of S0 is exactly what would be expected from the given distribution without any deviation. Therefore, the deviation of xi from x0i, $i\in I$, leads to the deviation of $f_N\left( x\right) $ from f(x). Thus, from (14) we have
\begin{eqnarray}
{\rm Var}\{x_i\}&=&\frac 1{\rho ^2(x_i)}{\rm Var}\{f_N(x_i)\}\n...
 ...  &=&\frac{f(x_i)[1-f(x_i)]}{\rho
^2(x_i)N},\qquad\forall i\in I. \end{eqnarray}
(19)
The $1\sigma $ distribution function deviation test for the distribution of f(x) leads to
\begin{displaymath}
\left\vert x_i-x_{0i}\right\vert <\frac 1{\rho (x_i)}\sqrt{\frac{f(x_i)[1-f(x_i)]}N},\quad \forall i\in I, \end{displaymath} (20)
which can be interpreted as the $1\sigma $ random variable deviation for the distribution.

One can verify that condition (20) can also be obtained by applying Eqs. (12) and (15) together with condition (9).

In the following, we present several statements concluded from Definition 1, which might be useful for statistical analysis.

Statement 1. In the cases for which random variables vary continuously, a sample passing the $1\sigma $ distribution function deviation test must also pass any other statistical test which depends on the deviation of a sample from its expected value at the $1\sigma $ confidence level.

Proof. Assuming that the random variables vary continuously, take sample $S\equiv \{x_i\vert i\in I\}$, $I\equiv \{i\vert 1\leq i\leq N\}$. Let a statistical function T of the random variables be


T=T(x1,x2,......,xN).

(21)

Then
\begin{eqnarray}
{\rm Var}\{T\}&=&\sum_{i=1}^N\left(\frac{\partial T}{\partial x...
 ... T}{\partial x_i}\right)^2\frac{f(x_i)[1-f(x_i)]}{\rho
^2(x_i)N}. \end{eqnarray}
(22)
Assuming sample S pass the $1\sigma $ distribution function deviation test, condition (20) is then satisfied. Therefore,

\begin{displaymath}
\left\vert T(x_1,x_2,......,x_N)-T(x_{01},x_{02},......,x_{0N})\right\vert =\end{displaymath}


\begin{displaymath}
\sqrt{\sum_{i=1}^N(\frac{\partial T}{\partial x_i})^2(x_i-x_{0i})^2}<\sqrt{{\rm Var}\{T\}}. \end{displaymath} (23)
This completes the proof.

Statement 2. If the random variables vary continuously, a sample passing the $1\sigma $ distribution function deviation test must also pass any mean-value test at the $1\sigma $ confidence level.

This statement is obvious according to Statement 1, as the mean value of any statistical function of a sample is also a statistical function of the random variables.

Statement 3. A sample passing the $1\sigma $ distribution function deviation test also passes the Kolmogorov-Smirnov test at the $1\sigma $confidence level.

Proof. Let sample $S\equiv \{x_i\vert i\in I\}$, $I\equiv \{i\vert 1\leq i\leq N\}$, pass the $1\sigma $ distribution function deviation test. According to (9), the following relation is satisfied:
\begin{displaymath}
\sqrt{N}\left\vert f_N(x)-f(x)\right\vert <{\sqrt{{f(x)[1-f(x)]}}.} \end{displaymath} (24)
However,
\begin{displaymath}
\max \left\{ {\sqrt{{f(x)[1-f(x)]}}}\right\} =0.5, \end{displaymath} (25)
since $0\leq f(x)\leq 1$. Therefore,
\begin{displaymath}
\sqrt{N}\sup _{-\infty <x<\infty }\left\vert f_N(x)-f(x)\right\vert <0.5{.} \end{displaymath} (26)
When N is large enough, the distribution of fN(x) at any given x is near Gaussian. The confidence level of $1\sigma $ is 0.683. Let $L(\lambda )=0.683$, where
\begin{displaymath}
L(\lambda )=1-2\sum_{j=1}^\infty (-1)^{j-1}\exp (-2j^2\lambda ^2). \end{displaymath} (27)
This gives $\lambda =0.96$. From (26) we find that
\begin{displaymath}
\sqrt{N}\sup _{-\infty <x<\infty }\left\vert f_N(x)-f(x)\right\vert <0.96{.} \end{displaymath} (28)
Thus, the sample passes the Kolmogorov-Smirnov test at the $1\sigma $ confidence level.

A comparison of Eqs. (26) and (28) shows that if a sample passes the $1\sigma $ distribution function deviation test, then it passes it far better than the K-S test at the $1\sigma $ confidence level, since $0.5\ll 0.96$.


next previous
Up: A statistical method for

Copyright The European Southern Observatory (ESO)