Parameters
Sampled parameters (e.g., from the prior distribution) are often stored as numobs/getobs) is supported.
It can sometimes be helpful to wrap the parameters in a user-defined type that also stores expensive intermediate objects needed for simulating data (e.g., Cholesky factors). The user-defined type should be a subtype of AbstractParameterSet, whose only requirement is a field θ that stores parameters.
For convenience, parameters can be stored with named dimensions; see, for example, NamedMatrix.
NeuralEstimators.AbstractParameterSet Type
AbstractParameterSetAn abstract supertype for user-defined types that store parameters and any auxiliary objects needed for data simulation.
The user-defined type must have a field θ that stores the parameters. Typically, θ is a numobs/getobs is supported. There are no other requirements.
The number of parameter instances can be retrieved with numobs, and the size of θ can be inspected with size.
Subtypes of AbstractParameterSet support indexing via Base.getindex, with any batchable fields subsetted accordingly and all other fields left unchanged. To modify this default behaviour, provide a specific Base.getindex method for your concrete subtype.
Examples
struct Parameters <: AbstractParameterSet
θ
# auxiliary objects needed for data simulation
end
θ = randn(2, 100)
parameters = Parameters(θ)
numobs(parameters) # 100
size(parameters) # (2, 100)
parameters[1:10] # subset of 10 parameter vectorsNamedArrays.NamedMatrix Type
NamedMatrix(; kwargs...)Returns a NamedArray with named rows (parameters) and indexed columns (samples).
Examples
NamedMatrix(μ = randn(3), σ = rand(3))Data
Simulated data sets are stored as mini-batches in a format amenable to the chosen neural-network architecture; the only requirement is compatibility with numobs/getobs. For example, when constructing an estimator from data collected over a two-dimensional grid, one may use a CNN, with each data set stored in the final dimension of a four-dimensional array.
Precomputed (expert) summary statistics can be incorporated by wrapping the simulated data and summary statistics in a DataAndSummaries object.
A vector of data sets (as used by DeepSet) can optionally be concatenated into a single array with PackedReplicates, so that a minibatch is one array rather than one array per data set.
NeuralEstimators.DataAndSummaries Type
DataAndSummaries(Z, S)A container that couples raw data Z (stored in a format amenable to the chosen neural-network architecture) with precomputed expert summary statistics S (a matrix whose columns are the summary statistics for each corresponding element of Z).
Passing a DataAndSummaries to any neural estimator causes the summary network to be applied to Z, with the resulting learned summary statistics concatenated with S before being passed to the inference network.
See also summarystatistics.
Examples
using NeuralEstimators
using Statistics: mean, var
# Simulate data: Z|μ,σ ~ N(μ, σ²)
n, m, K = 1, 50, 500
θ = rand(2, K)
Z = [θ[1, k] .+ θ[2, k] .* randn(n, m) for k in 1:K]
# Precompute expert summary statistics (e.g., sample mean and variance)
S = hcat([vcat(mean(z), var(z)) for z in Z]...)
# Package into a DataAndSummaries object
DataAndSummaries(Z, S)NeuralEstimators.PackedReplicates Type
PackedReplicates(Z::V; max_sample_size = nothing) where V <: AbstractVector{A} where A <: AbstractArrayA container that concatenates a vector of data sets into a single array, storing the original sample sizes alongside the packed data. Intended to be used with DeepSet.
Each element of Z is one data set, with exchangeable replicates stored in the last dimension. By default the packed data has final dimension of size sum(sample_sizes), where sample_sizes[i] is the number of replicates in the ith data set.
When max_sample_size is set, each data set is padded along its last dimension to that length before packing, so the packed data has final dimension max_sample_size * length(Z). A binary mask of size (max_sample_size, length(Z)) records which slots are real replicates. This fixed layout is required when training a DeepSet with Reactant on data sets of varying sample size.
Examples
using NeuralEstimators
# Original data
n = 2 # dimension of each data replicate
Z = [rand(Float32, n, m) for m in (3, 5, 4)]
# Packed data
P = PackedReplicates(Z) # data size (n, 12), sample_sizes == [3, 5, 4]
P[1:2] # first two data sets
# Fixed shape for Reactant
P = PackedReplicates(Z; max_sample_size = 5)