quickstats
Quickly compute simple statistics
Loading...
Searching...
No Matches
quickstats Namespace Reference

Quickly compute simple statistics. More...

Classes

struct  MadOptions
 Options for mad() and mad_with_infinities(). More...
 
struct  MedianOptions
 Options for median(). More...
 
class  MultipleQuantilesFixedNumber
 Calculate multiple quantiles from a fixed number of elements. More...
 
class  MultipleQuantilesVariableNumber
 Calculate multiple quantiles from a variable number of elements. More...
 
struct  MultipleQuantilesVariableNumberOptions
 Options for MultipleQuantilesVariableNumber. More...
 
struct  PairwiseSumOptions
 Options for pairwise_sum() and pairwise_sum_abstract(). More...
 
struct  PairwiseSumWorkspace
 Re-usable workspace for pairwise_sum() and pairwise_sum_abstract(). More...
 
struct  RssOptions
 Options for rss(). More...
 
struct  RssResult
 Result of rss(). More...
 
struct  RssWorkspace
 Re-usable workspace for rss(). More...
 
class  SingleQuantileFixedNumber
 Calculate a single quantile from a fixed number of elements. More...
 
class  SingleQuantileVariableNumber
 Calculate a quantile from a variable number of elements. More...
 
struct  SingleQuantileVariableNumberOptions
 Options for SingleQuantileVariableNumber. More...
 

Functions

template<typename Output_ = double, typename Input_ >
Output_ mad (const std::size_t num_total, Input_ *const ptr, const Input_ median, const MadOptions< Output_ > &options)
 
template<typename Output_ = double, typename Input_ >
Output_ mad (const std::size_t num_total, const std::size_t num_non_zero, Input_ *const values, const Input_ median, const MadOptions< Output_ > &options)
 
template<typename Float_ >
Float_ scale_mad_to_sd (const Float_ x)
 
template<typename Output_ = double, typename Input_ >
Output_ median (const std::size_t num_total, Input_ *const ptr, const MedianOptions< Output_ > &options)
 
template<typename Output_ = double, typename Input_ >
Output_ median (const std::size_t num_total, const std::size_t num_non_zero, Input_ *const values, const MedianOptions< Output_ > &options)
 
template<std::size_t accumulators_ = 4, class Input_ , typename Output_ >
Output_ pairwise_sum_abstract (const std::size_t num_total, Input_ input, PairwiseSumWorkspace< Output_ > &work, const PairwiseSumOptions &options)
 
template<std::size_t accumulators_ = 4, typename Input_ , typename Output_ >
Output_ pairwise_sum (const std::size_t num_total, const Input_ *const ptr, PairwiseSumWorkspace< Output_ > &work, const PairwiseSumOptions &options)
 
template<std::size_t accumulators_ = 4, typename Input_ , typename Output_ >
RssResult< Output_ > rss (const std::size_t num_total, const std::size_t num_non_zero, const Input_ *const ptr, RssWorkspace< Output_ > &work, const RssOptions< Output_ > &options)
 
template<std::size_t accumulators_ = 4, typename Input_ , typename Output_ >
RssResult< Output_ > rss (const std::size_t num_total, const Input_ *const ptr, RssWorkspace< Output_ > &work, const RssOptions< Output_ > &options)
 
template<typename Output_ = double, typename Input_ , typename Count_ >
void update_rss (Output_ &mean, Output_ &rss, const Input_ value, const Count_ num_total)
 
template<typename Output_ = double, typename Count_ >
void update_rss_with_zeros_unsafe (Output_ &mean, Output_ &rss, const Count_ num_zeros, const Count_ num_total)
 
template<typename Output_ = double, typename Count_ >
void update_rss_with_zeros (Output_ &mean, Output_ &rss, const Count_ num_zeros, const Count_ num_total)
 
template<typename Count_ , typename Float_ >
Float_ recenter_rss_unsafe (const Count_ num_total, const Float_ old_rss, const Float_ old_mean, const Float_ new_mean)
 
template<typename Count_ , typename Float_ >
Float_ recenter_rss (const Count_ num_total, const Float_ old_rss, const Float_ old_mean, const Float_ new_mean)
 
template<typename Input_ , class Skip_ >
std::size_t skip_values (const std::size_t num_total, Input_ *const ptr, Skip_ skip)
 
template<typename Value_ >
constexpr Value_ nan_if_available_else_zero ()
 

Detailed Description

Quickly compute simple statistics.

Function Documentation

◆ mad() [1/2]

template<typename Output_ = double, typename Input_ >
Output_ quickstats::mad ( const std::size_t num_total,
Input_ *const ptr,
const Input_ median,
const MadOptions< Output_ > & options )

Compute the median absolute deviation (MAD) of an array of elements, given its median.

No consideration is given to special values like NaNs in the array. If these are to be skipped, consider using skip_values() before calling this function.

See also scale_mad_to_sd() if the MAD is to be used as an estimate of the standard deviation.

Template Parameters
Output_Floating-point type of the output value.
Input_Numeric type of the input values and median. This is generally expected to be floating-point, though it is also possible to use signed integers as long as their differences do not overflow.
Parameters
num_totalTotal number of elements from which to compute a MAD.
[in]ptrPointer to an array of length num_total. On output, the contents will contain the (possibly reordered) absolute deviations from the median.
medianMedian of the array at ptr, typically computed with median(). For medians that might be infinite, consider using mad_with_infinities() instead.
optionsFurther options.
Returns
MAD of the array. If num_total == 0, MadOptions::placeholder is returned. If median is infinite and MadOptions::difference_between_infinites_is_zero == false, MadOptions::placeholder is also returned.

◆ mad() [2/2]

template<typename Output_ = double, typename Input_ >
Output_ quickstats::mad ( const std::size_t num_total,
const std::size_t num_non_zero,
Input_ *const values,
const Input_ median,
const MadOptions< Output_ > & options )

Compute the median absolute deviation (MAD) of a sparse vector, given its median. This vector is assumed to have num_non_zero structural non-zeros and num_total - num_non_zero zeros.

No consideration is given to special values like NaNs in the array. If these are to be skipped, consider using skip_values() before calling this function.

See also scale_mad_to_sd() if the MAD is to be used as an estimate of the standard deviation.

Template Parameters
Output_Floating-point type of the output value.
Input_Numeric type of the input values and median. This is generally expected to be floating-point, though it is also possible to use signed integers as long as their differences do not overflow.
Parameters
num_totalTotal number of elements in the sparse vector.
num_non_zeroNumber of structural non-zeros in the sparse vector. This should be no greater than num_total. num_total - num_non_zero is the number of structural zeros.
[in]valuesPointer to the start of an array of length num_non_zero, containing the values of the structural non-zeros of the sparse vector. On output, the contents will contain the (possibly reordered) absolute deviations from the median.
medianMedian of the sparse vector at values, typically computed with median(). For medians that might be infinite, consider using mad_with_infinities() instead.
optionsFurther options.
Returns
MAD of the sparse vector. If num_total == 0, MadOptions::placeholder is returned. If median is infinite and MadOptions::difference_between_infinites_is_zero == false, MadOptions::placeholder is also returned.

◆ scale_mad_to_sd()

template<typename Float_ >
Float_ quickstats::scale_mad_to_sd ( const Float_ x)

Scale the median absolute deviation (MAD) so that its expected value for a normal distribution is equal to the standard deviation.

Template Parameters
Float_Floating-point type of the MAD.
Parameters
xMAD, typically computed from mad().
Returns
Scaled value of x.

◆ median() [1/2]

template<typename Output_ = double, typename Input_ >
Output_ quickstats::median ( const std::size_t num_total,
Input_ *const ptr,
const MedianOptions< Output_ > & options )

Compute the median of an array of elements.

No consideration is given to special values like NaNs in the array. If these are to be skipped, consider using skip_values() before calling this function.

Parameters
num_totalTotal number of elements from which to compute a median.
[in]ptrPointer to an array of length num_total. On output, the contents may be reordered.
optionsFurther options.
Template Parameters
Output_Floating-point type of the output value.
Input_Numeric type of the input values.
Returns
Median of values in [ptr, ptr + num_total), or MedianOptions::placeholder if num_total == 0.

◆ median() [2/2]

template<typename Output_ = double, typename Input_ >
Output_ quickstats::median ( const std::size_t num_total,
const std::size_t num_non_zero,
Input_ *const values,
const MedianOptions< Output_ > & options )

Compute the median of a sparse vector. This vector is assumed to have num_non_zero structural non-zeros and num_total - num_non_zero zeros.

No consideration is given to special values like NaNs in the array. If these are to be skipped, consider using skip_values() before calling this function.

Parameters
num_totalTotal number of elements in the sparse vector.
num_non_zeroNumber of structural non-zeros in the sparse vector. This should be no greater than num_total. num_total - num_non_zero is the number of structural zeros.
[in]valuesPointer to the start of an array of length num_non_zero, containing the values of the structural non-zeros of the sparse vector. On output, the contents may be reordered.
optionsFurther options.
Template Parameters
Output_Floating-point type of the output value.
Input_Numeric type of the input values.
Returns
Median of the sparse vector, or MedianOptions::placeholder if num_total == 0.

◆ pairwise_sum_abstract()

template<std::size_t accumulators_ = 4, class Input_ , typename Output_ >
Output_ quickstats::pairwise_sum_abstract ( const std::size_t num_total,
Input_ input,
PairwiseSumWorkspace< Output_ > & work,
const PairwiseSumOptions & options )

Perform pairwise summation on an abstract array of numeric elements. The array is recursively halved until each subarray is no greater than PairwiseSumOptions::max_sum_length. Elements in each subarray are summed in sequence, and the subtotals are then added together to obtain the total sum. Compared to naive summation, this reduces round-off error from floating-point imprecision with minimal loss of performance.

Template Parameters
accumulators_Number of accumulators to sum each subarray. This should be positive and should be no greater than PairwiseSumOptions::max_sum_length / 2. It is strongly recommended that this is set to a power of 2. Larger values improve instruction parallelization at the cost of increased binary size and potential register spills.
Input_Function that accepts a std::size_t and returns an input value (typically numeric).
Output_Numeric type of the sum.
Parameters
num_totalTotal number of observations.
inputFunction that accesses an abstract array of length num_total. Specifically, it accepts an integer index in [0, num_total) and returns an input value to be summed. Calls to input() may be vectorized so it is assumed that (i) there are no dependencies between calls and (ii) no exceptions are thrown.
workWorkspace that can be re-used across multiple pairwise_sum_abstract() calls.
optionsFurther options.
Returns
Sum of all input(i) values for i in [0, num_total).

◆ pairwise_sum()

template<std::size_t accumulators_ = 4, typename Input_ , typename Output_ >
Output_ quickstats::pairwise_sum ( const std::size_t num_total,
const Input_ *const ptr,
PairwiseSumWorkspace< Output_ > & work,
const PairwiseSumOptions & options )

Perform pairwise summation on an array, see pairwise_sum_abstract() for more details.

Template Parameters
accumulators_Maximum number of elements to sum in sequence, see pairwise_sum_abstract() for details.
Input_Numeric type of the input data.
Output_Numeric type of the sum.
Parameters
num_totalTotal number of observations.
[in]ptrPointer to an array of length num_total, containing the elements to be summed.
workWorkspace that can be re-used across multiple pairwise_sum() calls.
optionsFurther options.
Returns
Sum of all elements in [ptr, ptr + num_total).

◆ rss() [1/2]

template<std::size_t accumulators_ = 4, typename Input_ , typename Output_ >
RssResult< Output_ > quickstats::rss ( const std::size_t num_total,
const std::size_t num_non_zero,
const Input_ *const ptr,
RssWorkspace< Output_ > & work,
const RssOptions< Output_ > & options )

Compute the residual sum of squares from a sparse vector. This uses the standard two-pass algorithm with naive accumulation of the sum of squared differences; thus, it is best used with a sufficiently high-precision Output_ like double.

No consideration is given to special values like NaNs in the values of the structural non-zeros. If these are to be skipped, consider using skip_values() before calling this method.

Template Parameters
accumulators_Number of accumulators, see pairwise_sum() for details.
Input_Numeric type of the input values.
Output_Floating-point type of the output data.
Parameters
num_totalTotal number of elements in the sparse vector.
num_non_zeroNumber of structural non-zeros in the sparse vector. This should be no greater than num_total. num_total - num_non_zero is the number of structural zeros.
[in]ptrPointer to an array of length num_non_zero, containing the values of the structural non-zeros in the sparse vector.
workWorkspace that can be re-used across multiple rss() calls.
optionsFurther options.
Returns
The sample mean and residual sum of squares of the sparse vector.

◆ rss() [2/2]

template<std::size_t accumulators_ = 4, typename Input_ , typename Output_ >
RssResult< Output_ > quickstats::rss ( const std::size_t num_total,
const Input_ *const ptr,
RssWorkspace< Output_ > & work,
const RssOptions< Output_ > & options )

Compute the residual sum of squares from the mean from a dense array. This uses the standard two-pass algorithm with naive accumulation of the sum of squared differences; thus, it is best used with a sufficiently high-precision Output_ like double.

No consideration is given to special values like NaNs in the dense array. If these are to be skipped, consider using skip_values() before calling this method.

Template Parameters
accumulators_Number of accumulators, see pairwise_sum() for details.
Input_Numeric type of the input values.
Output_Floating-point type of the output data.
Parameters
num_totalTotal number of elements in the array.
[in]ptrPointer to an array of length num_total.
workWorkspace that can be re-used across multiple rss() calls.
optionsFurther options.
Returns
The sample mean and residual sum of squares of the array.

◆ update_rss()

template<typename Output_ = double, typename Input_ , typename Count_ >
void quickstats::update_rss ( Output_ & mean,
Output_ & rss,
const Input_ value,
const Count_ num_total )

Update the mean and RSS by adding a new value using Welford's method.

Template Parameters
Output_Floating-point type of the output statistics.
Input_Numeric type of the inut value.
Count_Integer type of the number of values.
Parameters
meanOn input, the mean of previous values. If no previous values were provided, this should be set to zero. On output, the updated mean after including the latest value.
rssOn input, the RSS of previous values. If no previous values were provided, this should be set to zero. On output, the updated RSS after including the latest value.
valueNew value to update the mean/RSS.
num_totalNumber of values used to compute the updated mean/RSS, after adding the latest value. This should always be positive.

◆ update_rss_with_zeros_unsafe()

template<typename Output_ = double, typename Count_ >
void quickstats::update_rss_with_zeros_unsafe ( Output_ & mean,
Output_ & rss,
const Count_ num_zeros,
const Count_ num_total )

Update the mean and RSS by adding any number of zeros using Welford's method. This assumes that num_total > 0; if this cannot be guaranteed, use update_rss_with_zeros() instead.

Template Parameters
Output_Floating-point type of the output statistics.
Count_Integer type of the number of values.
Parameters
meanOn input, the mean of previous values. If no previous values were provided, this should be set to zero. On output, the updated mean after including the zeros.
rssOn input, the RSS of previous values. If no previous values were provided, this should be set to zero. On output, the updated RSS after including the zeros.
num_zerosNumber of zeros to be added. This may be zero.
num_totalNumber of values used to compute the updated mean/RSS, after adding the specified number of zeros. This should be positive and no less than num_total.

◆ update_rss_with_zeros()

template<typename Output_ = double, typename Count_ >
void quickstats::update_rss_with_zeros ( Output_ & mean,
Output_ & rss,
const Count_ num_zeros,
const Count_ num_total )

Update the mean and RSS with any number of zeros using Welford's method. This is a slightly slower version of update_rss_with_zeros_unsafe() that handles num_total == 0.

Template Parameters
Output_Floating-point type of the output statistics.
Count_Integer type of the number of values.
Parameters
meanOn input, the mean of previous values. If no previous values were provided, this should be set to zero. On output, the updated mean after including the zeros.
rssOn input, the RSS of previous values. If no previous values were provided, this should be set to zero. On output, the updated RSS after including the zeros.
num_zerosNumber of zero values to be added. This may be zero.
num_totalNumber of values used to compute the updated mean/RSS, after adding any number of zeros. This may be zero but should be no less than num_total.

◆ recenter_rss_unsafe()

template<typename Count_ , typename Float_ >
Float_ quickstats::recenter_rss_unsafe ( const Count_ num_total,
const Float_ old_rss,
const Float_ old_mean,
const Float_ new_mean )

Recenter the residual sum of squares, i.e., sum of squares from a different mean. This is typically used to combine RSS values from different subsets of the data, e.g., when splitting calculations across cores. In such cases, the global mean across the entire dataset should be used as the new mean, and the sum of the recentered RSS values will be the RSS of the entire dataset. (This approach is more numerically stable than computing the sum of squared observations and then computing the difference with the squared mean.)

This function is considered "unsafe" as it assumes that old_mean is zero when num_total == 0. In many cases, old_mean will be NaN when num_total == 0 due to division by zero. If num_total > 0 or old_mean == 0 cannot be guaranteed, consider using recenter_rss() instead.

Template Parameters
Count_Integer type of the number of values.
Float_Floating-point type of the various statistics.
Parameters
num_totalTotal number of values used to compute the RSS.
old_rssThe old value of the RSS.
old_meanThe old mean used to compute the RSS.
new_meanThe new mean.
Returns
The recentered RSS.

◆ recenter_rss()

template<typename Count_ , typename Float_ >
Float_ quickstats::recenter_rss ( const Count_ num_total,
const Float_ old_rss,
const Float_ old_mean,
const Float_ new_mean )

Recenter the residual sum of squares, i.e., sum of squares from a different mean. This is a safer version of recenter_rss_unsafe() that correctly handles num_total == 0 and an old_mean of NaN, at the cost of some performance.

Template Parameters
Count_Integer type of the number of values.
Float_Floating-point type of the various statistics.
Parameters
num_totalTotal number of values used to compute the RSS. This should be non-negative.
old_rssThe old value of the RSS.
old_meanThe old mean used to compute the RSS. If num_total == 0, this is ignored and may be set to NaN; this NaN will not propagate to the output.
new_meanThe new mean.
Returns
The recentered RSS, or old_rss (which should be zero) if num_total == 0.

◆ skip_values()

template<typename Input_ , class Skip_ >
std::size_t quickstats::skip_values ( const std::size_t num_total,
Input_ *const ptr,
Skip_ skip )
Template Parameters
Input_Numeric type of the input.
Skip_Function defining what elements to skip.
Parameters
num_totalTotal number of elements.
[in,out]ptrPointer to an array of length num_total, containing the input elements. On output, the first \(n\) entries will contain the unskipped elements, in the same order as provided in the input.
skipFunction that accepts the index of an element in ptr (as a std::size_t) and the value of the element (as an Input_), and returns a boolean indicating whether this element should be skipped.
Returns
The number of unskipped elements \(n\).

◆ nan_if_available_else_zero()

template<typename Value_ >
Value_ quickstats::nan_if_available_else_zero ( )
constexpr
Template Parameters
Value_Some numeric type.
Returns
NaN if supported by Value_, otherwise zero.