Data Package (ria_toolkit_oss.data)

The Data package contains abstract data types tailored for radio machine learning, such as Recording, as well as the abstract interfaces for the radio dataset and radio dataset builder framework.

class ria_toolkit_oss.data.Annotation(sample_start, sample_count, freq_lower_edge, freq_upper_edge, label='', comment='', detail=None)[source]

Bases: object

Signal annotations are labels or additional information associated with specific data points or segments within a signal. These annotations could be used for tasks like supervised learning, where the goal is to train a model to recognize patterns or characteristics in the signal associated with these annotations.

Annotations can be used to label interesting points in your recording.

Parameters:
  • sample_start (int) – The index of the starting sample of the annotation.

  • sample_count (int) – The index of the ending sample of the annotation, inclusive.

  • freq_lower_edge (float) – The lower frequency of the annotation.

  • freq_upper_edge (float) – The upper frequency of the annotation.

  • label (str, optional) – The label that will be displayed with the bounding box in compatible viewers including IQEngine. Defaults to an emtpy string.

  • comment (str, optional) – A human-readable comment. Defaults to an empty string.

  • detail (dict, optional) – A dictionary of user defined annotation-specific metadata. Defaults to None.

is_valid()[source]

Verify sample_count > 0 and the freq_lower_edge < freq_upper_edge.

Returns:

True if valid, False if not.

overlap(other)[source]

Quantify how much the bounding box in this annotation overlaps with another annotation.

Parameters:

other (Annotation) – The other annotation.

Returns:

The area of the overlap in samples*frequency, or 0 if they do not overlap.

area()[source]

The ‘area’ of the bounding box, samples*frequency. Useful to quantify annotation size.

Returns:

sample length multiplied by bandwidth.

to_sigmf_format()[source]

Returns a JSON dictionary representation, formatted for saving in a .sigmf-meta file.

class ria_toolkit_oss.data.Recording(data, metadata=None, dtype=None, timestamp=None, annotations=None)[source]

Bases: object

Tape of complex IQ (in-phase and quadrature) samples with associated metadata and annotations.

Recording data is a complex array of shape C x N, where C is the number of channels and N is the number of samples in each channel.

Metadata is stored in a dictionary of key value pairs, to include information such as sample_rate and center_frequency.

Annotations are a list of Annotation, defining bounding boxes in time and frequency with labels and metadata.

Here, signal data is represented as a NumPy array. This class is then extended in the RIA Backends to provide support for different data structures, such as Tensors.

Recordings are long-form tapes can be obtained either from a software-defined radio (SDR) or generated synthetically. Then, machine learning datasets are curated from collection of recordings by segmenting these longer-form tapes into shorter units called slices.

All recordings are assigned a unique 64-character recording ID, rec_id. If this field is missing from the provided metadata, a new ID will be generated upon object instantiation.

Parameters:
  • data (array_like) – Signal data as a tape IQ samples, either C x N complex, where C is the number of channels and N is number of samples in the signal. If data is a one-dimensional array of complex samples with length N, it will be reshaped to a two-dimensional array with dimensions 1 x N.

  • metadata (dict, optional) – Additional information associated with the recording.

  • annotations (list of Annotations, optional) – A collection of Annotation objects defining bounding boxes.

  • dtype (float or int, optional) – Explicitly specify the data-type of the complex samples. Must be a complex NumPy type, such as numpy.complex64 or numpy.complex128. Default is None, in which case the type is determined implicitly. If data is a NumPy array, the Recording will use the dtype of data directly without any conversion.

  • timestamp – The timestamp when the recording data was generated. If provided, it should be a float or integer representing the time in seconds since epoch (e.g., time.time()). Only used if the timestamp field is not present in the provided metadata.

Raises:
  • ValueError – If data is not complex 1xN or CxN.

  • ValueError – If metadata is not a python dict.

  • ValueError – If metadata is not json serializable.

  • ValueError – If annotations is not a list of valid annotation objects.

Examples:

>>> import numpy
>>> from ria_toolkit_oss.data import Recording, Annotation
>>> # Create an array of complex samples, just 1s in this case.
>>> samples = numpy.ones(10000, dtype=numpy.complex64)
>>> # Create a dictionary of relevant metadata.
>>> sample_rate = 1e6
>>> center_frequency = 2.44e9
>>> metadata = {
...     "sample_rate": sample_rate,
...     "center_frequency": center_frequency,
...     "author": "me",
... }
>>> # Create an annotation for the annotations list.
>>> annotations = [
...     Annotation(
...         sample_start=0,
...         sample_count=1000,
...         freq_lower_edge=center_frequency - (sample_rate / 2),
...         freq_upper_edge=center_frequency + (sample_rate / 2),
...         label="example",
...     )
... ]
>>> # Store samples, metadata, and annotations together in a convenient object.
>>> recording = Recording(data=samples, metadata=metadata, annotations=annotations)
>>> print(recording.metadata)
{'sample_rate': 1000000.0, 'center_frequency': 2440000000.0, 'author': 'me'}
>>> print(recording.annotations[0].label)
'example'
property data
Returns:

Recording data, as a complex array.

Type:

ndarray

Note

For recordings with more than 1,024 samples, this property returns a read-only view of the data.

Note

To access specific samples, consider indexing the object directly with rec[c, n].

property metadata
Returns:

Dictionary of recording metadata.

Type:

dict

property annotations
Returns:

List of recording annotations

Type:

list of Annotation objects

property shape
Returns:

The shape of the data array.

Type:

tuple of ints

property n_chan
Returns:

The number of channels in the recording.

Type:

int

property rec_id
Returns:

Recording ID.

Type:

str

property dtype
Returns:

Data-type of the data array’s elements.

Type:

numpy dtype object

property timestamp
Returns:

Recording timestamp (time in seconds since epoch).

Type:

float or int

property sample_rate
Returns:

Sample rate of the recording, or None if ‘sample_rate’ is not in metadata.

Type:

str

astype(dtype)[source]

Copy of the recording, data cast to a specified type.

Parameters:

dtype (NumPy data type, optional) – Data-type to which the array is cast. Must be a complex scalar type, such as numpy.complex64 or numpy.complex128.

Returns:

A new recording with the same metadata and data, with dtype.

Examples:

Todo

Usage examples coming soon!

add_to_metadata(key, value)[source]

Add a new key-value pair to the recording metadata.

Parameters:
  • key (str) – New metadata key, must be snake_case.

  • value (any) – Corresponding metadata value.

Raises:
  • ValueError – If key is already in metadata or if key is not a valid metadata key.

  • ValueError – If value is not JSON serializable.

Returns:

None.

Examples:

Create a recording and add metadata:

>>> import numpy
>>> from ria_toolkit_oss.data import Recording
>>>
>>> samples = numpy.ones(10000, dtype=numpy.complex64)
>>> metadata = {
>>>     "sample_rate": 1e6,
>>>     "center_frequency": 2.44e9,
>>> }
>>>
>>> recording = Recording(data=samples, metadata=metadata)
>>> print(recording.metadata)
{'sample_rate': 1000000.0,
'center_frequency': 2440000000.0,
'timestamp': 17369...,
'rec_id': 'fda0f41...'}
>>>
>>> recording.add_to_metadata(key="author", value="me")
>>> print(recording.metadata)
{'sample_rate': 1000000.0,
'center_frequency': 2440000000.0,
'author': 'me',
'timestamp': 17369...,
'rec_id': 'fda0f41...'}
update_metadata(key, value)[source]

Update the value of an existing metadata key, or add the key value pair if it does not already exist.

Parameters:
  • key (str) – Existing metadata key.

  • value (any) – New value to enter at key.

Raises:
Returns:

None.

Examples:

Create a recording and update metadata:

>>> import numpy
>>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64)
>>> metadata = {
>>>     "sample_rate": 1e6,
>>>     "center_frequency": 2.44e9,
>>>     "author": "me"
>>> }
>>> recording = Recording(data=samples, metadata=metadata)
>>> print(recording.metadata)
{'sample_rate': 1000000.0,
'center_frequency': 2440000000.0,
'author': "me",
'timestamp': 17369...
'rec_id': 'fda0f41...'}
>>> recording.update_metadata(key="author", value=you")
>>> print(recording.metadata)
{'sample_rate': 1000000.0,
'center_frequency': 2440000000.0,
'author': "you",
'timestamp': 17369...
'rec_id': 'fda0f41...'}
remove_from_metadata(key)[source]

Remove a key from the recording metadata. Does not remove key if it is protected.

Parameters:

key (str) – The key to remove.

Raises:

ValueError – If key is protected.

Returns:

None.

Examples:

Create a recording and add metadata:

>>> import numpy
>>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64)
>>> metadata = {
...     "sample_rate": 1e6,
...     "center_frequency": 2.44e9,
... }
>>> recording = Recording(data=samples, metadata=metadata)
>>> print(recording.metadata)
{'sample_rate': 1000000.0,
'center_frequency': 2440000000.0,
'timestamp': 17369...,  # Example value
'rec_id': 'fda0f41...'}  # Example value
>>> recording.add_to_metadata(key="author", value="me")
>>> print(recording.metadata)
{'sample_rate': 1000000.0,
'center_frequency': 2440000000.0,
'author': 'me',
'timestamp': 17369...,  # Example value
'rec_id': 'fda0f41...'}  # Example value
view(output_path='images/signal.png', **kwargs)[source]

Create a plot of various signal visualizations as a PNG image.

Parameters:
  • output_path (str, optional) – The output image path. Defaults to “images/signal.png”.

  • kwargs – Keyword arguments passed on to ria_toolkit_oss.view.view_sig.

Type:

dict of keyword arguments

Examples:

Create a recording and view it as a plot in a .png image:

>>> import numpy
>>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64)
>>> metadata = {
>>>     "sample_rate": 1e6,
>>>     "center_frequency": 2.44e9,
>>> }
>>> recording = Recording(data=samples, metadata=metadata)
>>> recording.view()
simple_view(**kwargs)[source]

Create a plot of various signal visualizations as a PNG or SVG image.

Parameters:

kwargs – Keyword arguments passed on to ria_toolkit_oss.view.view_signal_simple.view_simple_sig.

Type:

dict of keyword arguments

Examples:

Create a recording and view it as a plot in a .png image:

>>> import numpy
>>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64)
>>> metadata = {
>>>     "sample_rate": 1e6,
>>>     "center_frequency": 2.44e9,
>>> }
>>> recording = Recording(data=samples, metadata=metadata)
>>> recording.simple_view()
to_sigmf(filename=None, path=None, overwrite=False)[source]

Write recording to a set of SigMF files.

The SigMF io format is defined by the SigMF Specification Project

Parameters:
  • recording (Recording) – The recording to be written to file.

  • filename (PathLike or str, optional) – The name of the file where the recording is to be saved. Defaults to auto generated filename.

  • path (PathLike or str, optional) – The directory path to where the recording is to be saved. Defaults to recordings/.

Raises:

IOError – If there is an issue encountered during the file writing process.

Returns:

None

to_npy(filename=None, path=None, overwrite=False)[source]

Write recording to .npy binary file.

Parameters:
  • filename (PathLike or str, optional) – The name of the file where the recording is to be saved. Defaults to auto generated filename.

  • path (PathLike or str, optional) – The directory path to where the recording is to be saved. Defaults to recordings/.

Raises:

IOError – If there is an issue encountered during the file writing process.

Returns:

Path where the file was saved.

Return type:

str

Examples:

Create a recording and save it to a .npy file:

>>> import numpy
>>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64)
>>> metadata = {
>>>     "sample_rate": 1e6,
>>>     "center_frequency": 2.44e9,
>>> }
>>> recording = Recording(data=samples, metadata=metadata)
>>> recording.to_npy()
to_wav(filename=None, path=None, target_sample_rate=48000, bits_per_sample=32, overwrite=False)[source]

Write recording to WAV file with embedded YAML metadata.

WAV format uses stereo audio with I (in-phase) in left channel and Q (quadrature) in right channel. Metadata is stored in standard LIST INFO chunks with RF-specific metadata encoded as YAML in the ICMT (comment) field for human readability.

Parameters:
  • filename (PathLike or str, optional) – The name of the file where the recording is to be saved. Defaults to auto generated filename.

  • path (PathLike or str, optional) – The directory path to where the recording is to be saved. Defaults to recordings/.

  • target_sample_rate (int, optional) – Sample rate stored in the WAV header when no sample_rate metadata is present. IQ samples are written without decimation or interpolation. Default is 48000 Hz.

  • bits_per_sample (int, optional) – Bits per sample (32 for float32, 16 for int16). Default is 32.

  • overwrite (bool, optional) – Whether to overwrite existing files. Default is False.

Raises:

IOError – If there is an issue encountered during the file writing process.

Returns:

Path where the file was saved.

Return type:

str

Examples:

Create a recording and save it to a .wav file:

>>> import numpy
>>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.exp(1j * 2 * numpy.pi * 0.1 * numpy.arange(10000))
>>> metadata = {"sample_rate": 1e6, "center_frequency": 915e6}
>>> recording = Recording(data=samples, metadata=metadata)
>>> recording.to_wav()
to_blue(filename=None, path=None, data_format='CI', overwrite=False)[source]

Write recording to MIDAS Blue file format.

MIDAS Blue is a legacy RF file format with a 512-byte binary header. Commonly used with X-Midas and other RF/radar signal processing tools.

Parameters:
  • filename (PathLike or str, optional) – The name of the file where the recording is to be saved. Defaults to auto generated filename.

  • path (PathLike or str, optional) – The directory path to where the recording is to be saved. Defaults to recordings/.

  • data_format (str, optional) – Format code (default ‘CI’ = complex int16). Common formats: ‘CI’ (complex int16), ‘CF’ (complex float32), ‘CD’ (complex float64). Integer formats require the IQ samples to already be scaled within [-1, 1).

  • overwrite (bool, optional) – Whether to overwrite existing files. Default is False.

Raises:

IOError – If there is an issue encountered during the file writing process.

Returns:

Path where the file was saved.

Return type:

str

Examples:

Create a recording and save it to a .blue file:

>>> import numpy
>>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64)
>>> metadata = {"sample_rate": 1e6, "center_frequency": 2.44e9}
>>> recording = Recording(data=samples, metadata=metadata)
>>> recording.to_blue()
trim(num_samples, start_sample=0)[source]

Trim Recording samples to a desired length, shifting annotations to maintain alignment.

Parameters:
  • start_sample (int, optional) – The start index of the desired trimmed recording. Defaults to 0.

  • num_samples (int) – The number of samples that the output trimmed recording will have.

Raises:
  • IndexError – If start_sample + num_samples is greater than the length of the recording.

  • IndexError – If sample_start < 0 or num_samples < 0.

Returns:

The trimmed Recording.

Return type:

Recording

Examples:

Create a recording and trim it:

>>> import numpy
>>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64)
>>> metadata = {
...     "sample_rate": 1e6,
...     "center_frequency": 2.44e9,
... }
>>> recording = Recording(data=samples, metadata=metadata)
>>> print(len(recording))
10000
>>> trimmed_recording = recording.trim(start_sample=1000, num_samples=1000)
>>> print(len(trimmed_recording))
1000
normalize()[source]

Scale the recording data, relative to its maximum value, so that the magnitude of the maximum sample is 1.

Returns:

Recording where the maximum sample amplitude is 1.

Return type:

Recording

Examples:

Create a recording with maximum amplitude 0.5 and normalize to a maximum amplitude of 1:

>>> import numpy
>>> from ria_toolkit_oss.data import Recording
>>> samples = numpy.ones(10000, dtype=numpy.complex64) * 0.5
>>> metadata = {
...     "sample_rate": 1e6,
...     "center_frequency": 2.44e9,
... }
>>> recording = Recording(data=samples, metadata=metadata)
>>> print(numpy.max(numpy.abs(recording.data)))
0.5
>>> normalized_recording = recording.normalize()
>>> print(numpy.max(numpy.abs(normalized_recording.data)))
1

Radio Dataset SubPackage

The Radio Dataset Subpackage defines the abstract interfaces and framework components for the management of machine learning datasets tailored for radio signal processing.

class ria_toolkit_oss.data.datasets.RadioDataset(source)[source]

Bases: ABC

A radio dataset is an iterable dataset designed for machine learning applications in radio signal processing and analysis. They are a structured collections of examples in a machine learning-ready format, with associated metadata.

This is an abstract interface defining common properties and behavior of radio datasets. Therefore, this class should not be instantiated directly. Instead, it should be subclassed to define specific interfaces for different types of radio datasets. For example, see ria_toolkit_oss.data.datasets.IQDataset, which is a radio dataset subclass tailored for tasks involving the processing of radio signals represented as IQ (In-phase and Quadrature) samples.

Parameters:

source (str or PathLike) – Path to the dataset source file. For more information on dataset source files and their format, see Intro to radio datasets.

property source
Returns:

Path to the dataset source file.

Type:

Path

property shape
Returns:

The shape of the dataset. The elements of the shape tuple give the lengths of the corresponding dataset dimensions.

Type:

tuple of ints

property data

Retrieve the data from the source file.

Note

Accessing this property reads all the data from the source file into memory as a NumPy array, which can consume significant amounts of memory and potentially degrade performance. Instead, use the RadioDataset class methods to process and manipulate the dataset source file. You can read individual examples into memory as NumPy arrays by indexing the dataset: RadioDataset[idx].

Returns:

The dataset examples as a single NumPy array.

Type:

ndarray

property metadata

Retrieve the metadata from the source file.

Note

Accessing this property reads all the metadata from the source file into memory as a Pandas DataFrame.

Returns:

The dataset metadata as a Pandas DataFrame.

Type:

pd.DataFrame

property labels

Retrieves the metadata labels from the dataset file.

Returns:

A list of metadata column headers.

Return type:

list of strings

Examples:

>>> awgn_builder = AWGN_Builder()
>>> awgn_builder.download_and_prepare()
>>> ds = awgn_builder.as_dataset(backend="pytorch")
>>> print(ds.labels)
['rec_id', 'modulation', 'snr']
abstract inspect()[source]

Todo

This method is not yet fully conceptualized. Likely, it will wrap some of the functionality in the Dataset Inspector package (dataset_manager.inspector) to produce an image or visualization. However, the Dataset Inspector package is not yet implemented.

abstract default_augmentations()[source]

Returns a list of default augmentations.

Returns:

A list of default augmentations.

Return type:

list of callable

augment(class_key, augmentations=None, level=1.0, target_size=None, classes_to_augment=None, inplace=False)[source]

Supplement the dataset with new examples by applying various transformations to the pre-existing examples in the dataset.

Todo

This method is currently under construction, and may produce unexpected results.

The process of supplementing a dataset to artificially increase the diversity of examples is called augmentation. Training on augmented data can enhance the generalization and robustness of deep machine learning models. For more information, see A Complete Guide to Data Augmentation.

Metadata for each new example will be identical to the metadata of the pre-existing example from which it was generated. The metadata will be extended to include an ‘augmentation’ column, populated with the string representation of the transform used.

Augmented data should only be used for model training, not for testing or validation.

Unless specified, augmentations are applied equally across classes, maintaining the original class distribution.

If target_size does not match the sum of the original class sizes scaled by an integer multiple, the class distribution is slightly adjusted to satisfy target_size.

Parameters:
  • class_key (str) – Class name used to augment from and calculate class distribution.

  • augmentations (callable or list of callables, optional) – A function or list of functions that take an example and return a transformed version. Defaults to default_augmentations().

  • level (float or list of floats, optional) –

    The extent of augmentation from 0.0 (none) to 1.0 (full). If classes_to_augment is specified, can be either:

    • A single float: All classes augmented evenly to this level.

    • A list of floats: Each element corresponds to the augmentation level target for the corresponding class.

  • target_size (int or list of ints, optional) –

    Target size of the augmented dataset. Overrides level if specified. If classes_to_augment is specified, can be either:

    • A single float: All classes are augmented proportional to their relative frequency until the dataset reaches target_size.

    • A list of floats: Each element corresponds to the target size for the corresponding class.

  • classes_to_augment (string or list of strings, optional) – List of metadata keys of classes to augment.

  • inplace (bool, optional) – If True, the augmentation is performed inplace and None is returned.

Raises:
  • ValueError – If level has any values not in the range (0,1].

  • ValueError – If target_size of dataset is already sufficed.

  • ValueError – If a class in classes_to_augment does not exist in class_key.

Returns:

The augmented dataset or None if inplace=True.

Return type:

RadioDataset or None

Examples:

>>> from ria.dataset_manager.builders import AWGN_Builder
>>> builder = AWGN_Builder()
>>> builder.download_and_prepare()
>>> ds = builder.as_dataset()
>>> ds.get_class_sizes(class_key='col')
{'a': 100, 'b': 500, 'c': 300}
>>> new_ds = ds.augment(class_key='col', classes_to_augment=['a', 'b'], target_size=1200)
>>> new_ds.get_class_sizes(class_key='col')
{'a': 150, 'b': 750, 'c': 300}
subsample(class_key, percentage, inplace=False)[source]

Reduces the number of examples in all classes of a dataset by randomly subsampling each class according to a specified percentage. This function reduces the number of examples per class to the specified percentage without affecting the overall class distribution.

Parameters:
  • class_key (str) – The name of the class to subsample.

  • percentage (float) – The percentage of the original class sizes to keep.

  • inplace (bool, optional) – If True, the operation modifies the existing source file directly and returns None. If False, the operation creates a new dataset object and corresponding source file, leaving the original dataset unchanged. Default is False.

Raises:

ValueError – If the target size of the class with the lowest frequency goes to 0.

Returns:

The subsampled dataset.

Return type:

RadioDataset or None

Examples:

>>> from ria.dataset_manager.builders import AWGN_Builder()
>>> builder = AWGN_Builder()
>>> builder.download_and_prepare()
>>> ds = builder.as_dataset()
>>> ds.get_class_sizes(class_key="col")
{a:100, b:200, c:300}
>>> new_ds = ds.subsample(percentage=0.80, class_key="col")
>>> new_ds.get_class_sizes(class_key="col")
{a:80, b:160, c:240}
resample(quantity_target, class_key, inplace=False)[source]

Adjusts an unsampled dataset by changing the number of examples per class to a user-specified quantity.

For each class:
  • If there are excess examples, it randomly subsamples the class to the quantity target.

  • If there are less examples, it randomly duplicates examples to reach the quantity target.

Parameters:
  • quantity_target (int) – The number of examples each class should have.

  • class_key (str) – The label of the class to resample.

  • inplace (bool, optional) – If True, the operation modifies the existing source file directly and returns None. If False, the operation creates a new dataset object and corresponding source file, leaving the original dataset unchanged. Default is False.

Returns:

The resampled dataset.

Return type:

RadioDataset or None

Examples:

>>> from ria.dataset_manager.builders import AWGN_Builder()
>>> builder = AWGN_Builder()
>>> builder.download_and_prepare()
>>> ds = builder.as_dataset()
>>> ds.get_class_sizes(class_key="col")
{a:100, b:200, c:300}
>>> new_ds = ds.resample(quantity_target=250, class_key="col")
>>> new_ds.get_class_sizes(class_key="col")
{a:250, b:250, c:250}
homogenize(class_key, example_limit=None, inplace=False)[source]
Discards excess samples by randomly subsampling all classes within a dataset that have more than a

user-specified limit of examples. If the user doesn’t specify a limit, the class the with the fewest examples is selected as the limit.

Parameters:
  • class_key (str) – The label of the class to homogenize.

  • example_limit (int, optional) – The class size limit to which all classes are subsampled. If not specified, the class with the fewest examples is used as the limit. Default is None.

  • inplace (bool, optional) – If True, the operation modifies the existing source file directly and returns None. If False, the operation creates a new dataset cbject and corresponding source file, leaving the original dataset unchanged. Default is False.

Returns:

The homogenized dataset.

Return type:

RadioDataset or None

Examples:

>>> from ria.dataset_manager.builders import AWGN_Builder()
>>> builder = AWGN_Builder()
>>> builder.download_and_prepare()
>>> ds = builder.as_dataset()
>>> ds.get_class_sizes(class_key="col")
{a:1000, b:5000, c:1500, d:900}
>>> new_ds = ds.homogenize(example_limit=1000, class_key="col")
>>> new_ds.get_class_sizes(class_key="col")
{a:1000, b:1000, c:1000, d:900}
>>> from ria.dataset_manager.builders import AWGN_Builder()
>>> builder = AWGN_Builder()
>>> builder.download_and_prepare()
>>> ds = builder.as_dataset()
>>> ds.get_class_sizes(class_key="col")
{a:1000, b:5000, c:1500, d:900}
>>> new_ds = ds.homogenize(class_key="col")
>>> new_ds.get_class_sizes(class_key="col")
{a:900, b:900, c:900, d:900}
drop_class(class_key, class_value, inplace=False)[source]

Removes an entire class from the dataset.

Parameters:
  • class_key (str) – Class that will have a value dropped from it. Example: ‘signal_type’

  • class_value (str) – Value of the class to be dropped. Example: ‘LTE’, ‘NR’

  • inplace (bool, optional) – If True, the operation modifies the existing source file directly and returns None. If False, the operation creates a new dataset cbject and corresponding source file, leaving the original dataset unchanged. Defaults to False.

Raises:

ValueError – If the entered class name does not exist in the dataset.

Returns:

The dataset without the removed class.

Return type:

RadioDataset or None

Examples:

>>> from ria.dataset_manager.builders import AWGN_Builder()
>>> builder = AWGN_Builder()
>>> builder.download_and_prepare()
>>> ds = builder.as_dataset()
>>> ds.get_class_sizes()
{a:100, b:500, c:300}
>>> new_ds = ds.drop_class('a')
>>> new_ds.get_class_sizes()
{b:500, c:300}
add_label(column_name, data, inplace=False)[source]

Add a new metadata label to the dataset.

Todo

This method is not yet implemented.

Parameters:
  • column_name – Name of the new metadata column header.

  • data – The contents of the new metadata column.

  • inplace (bool, optional) – If True, the label is added inplace and None is returned. Defaults to False.

Raises:

ValueError – If the length of data is not equal to the length of the dataset.

Returns:

The augmented dataset or None if inplace=True.

Return type:

RadioDataset or None

Examples:

Todo

Usage examples coming soon.

get_class_sizes(class_key)[source]

Returns a dictionary containing the sizes of each class in the dataset at the provided key.

Parameters:

class_key (str) – The class label.

Raises:

ValueError – If the specified key is not found in the dataset labels.

Returns:

A dictionary where each key is a distinct class label, and it’s value is the class size.

Return type:

A dictionary where the keys are strings and the values are integers

Examples:

>>> from ria.dataset_manager.builders import AWGN_Builder()
>>> spectrogram_sensing_builder = AWGN_Builder()
>>> spectrogram_sensing_builder.download_and_prepare()
>>> ds = spectrogram_sensing_builder.as_dataset(backend="pytorch")
>>> ds.get_class_sizes(class_key='signal_type')
{'LTE': 900, 'NR': 900, 'LTE_NR': 900}
delete_example(idx, inplace=False)[source]

Deletes an example and it’s corresponding metadata from the dataset.

Parameters:
  • idx (int) – The index of the example to be deleted.

  • inplace (bool, optional) – If True, the deletion is performed inplace and None is returned. Defaults to False.

Returns:

The new dataset or None if inplace=True.

Return type:

RadioDataset or None

Examples:

>>> from ria.dataset_manager.builders import AWGN_Builder()
>>> spectrogram_sensing_builder = AWGN_Builder()
>>> spectrogram_sensing_builder.download_and_prepare()
>>> ds = spectrogram_sensing_builder.as_dataset(backend="pytorch")
>>> len(ds)
2700
>>> ds = ds.delete_example(idx=34)
>>> len(ds)
2699
append(example, metadata)[source]

Append a single example to the end of the dataset. This operation is performed inplace.

Todo

This method is not yet implemented.

Parameters:
  • example (numpy.typing.ArrayLike) – The example to append.

  • metadata (dict) – The corresponding metadata dictionary.

Raises:

ValueError – If example does not the same shape and type as rest of the examples in the dataset.

Returns:

None.

Examples:

Todo

Usage examples coming soon.

join(ds)[source]

Join or merge together two radio datasets.

Todo

This method is not yet implemented.

  • Duplicate entries are not removed; they are included.

  • The examples are not shuffled; examples from ds are appended at the end.

  • Metadata will be expanded to contain all columns.

Parameters:

ds (bool) – The dataset to merge together with self. Examples from both datasets must have the same shape.

Returns:

The combined dataset.

Return type:

RadioDataset

Examples:

Todo

Usage examples coming soon.

filter(mask, inplace=False)[source]

Filter the dataset using the provided mask.

Todo

This method is not yet implemented.

Parameters:
  • mask (array_like) – A boolean mask. Where True, keep the corresponding examples. Where False, discard keep the corresponding examples. The filtering mask is often the result of applying a condition across the elements of the dataset.

  • inplace (bool, optional) – If True, the filter operation is performed inplace and None is returned. Defaults to False.

Returns:

The filtered dataset or None if inplace=True.

Return type:

RadioDataset or None

Examples:

Todo

Usage examples coming soon!

class ria_toolkit_oss.data.datasets.IQDataset(source)[source]

Bases: RadioDataset, ABC

An IQDataset is a RadioDataset tailored for machine learning tasks that involve processing radiofrequency (RF) signals represented as In-phase (I) and Quadrature (Q) samples.

For machine learning tasks that involve processing spectrograms, please use ria_toolkit_oss.data.datasets.SpectDataset instead.

This is an abstract interface defining common properties and behaviour of IQDatasets. Therefore, this class should not be instantiated directly. Instead, it is subclassed to define custom interfaces for specific machine learning backends.

Parameters:

source (str or PathLike) – Path to the dataset source file. For more information on dataset source files and their format, see Intro to radio datasets.

property shape
IQ datasets are M x C x N, where M is the number of examples, C is the number of channels, N is the length

of the signals.

Returns:

The shape of the dataset. The elements of the shape tuple give the lengths of the corresponding dataset dimensions.

Type:

tuple of ints

trim_examples(trim_length, keep='start', inplace=False)[source]

Trims all examples in a dataset to a desired length.

Parameters:
  • trim_length (int) – The desired length of the trimmed examples.

  • keep (str, optional) – Specifies the part of the example to keep. Defaults to “start”. The options are: - “start” - “end” - “middle” - “random”

  • inplace (bool) – If True, the operation modifies the existing source file directly and returns None. If False, the operation creates a new dataset cbject and corresponding source file, leaving the original dataset unchanged. Default is False.

Raises:
  • ValueError – If trim_length is greater than or equal to the length of the examples.

  • ValueError – If value of keep is not recognized.

  • ValueError – If specified trim length is invalid for middle index.

Returns:

The dataset that is composed of shorter examples.

Return type:

IQDataset

Examples:

>>> from ria.dataset_manager.builders import AWGN_Builder()
>>> builder = AWGN_Builder()
>>> builder.download_and_prepare()
>>> ds = builder.as_dataset()
>>> ds.shape
(5, 1, 3)
>>> new_ds = ds.trim_examples(2)
>>> new_ds.shape
(5, 1, 2)
split_examples(split_factor=None, example_length=None, inplace=False)[source]

If the current example length is not evenly divisible by the provided example_length, excess samples are discarded. Excess examples are always at the end of the slice. If the split factor results in non-integer example lengths for the new example chunks, it rounds down.

For example:

Requires either split_factor or example_length to be specified but not both. If both are provided, split factor will be used by default, and a warning will be raised.

Parameters:
  • split_factor (int, optional) – the number of new example chunks produced from each original example, defaults to None.

  • example_length (int, optional) – the example length of the new example chunks, defaults to None.

  • inplace (bool, optional) – If True, the operation modifies the existing source file directly and returns None. If False, the operation creates a new dataset cbject and corresponding source file, leaving the original dataset unchanged. Default is False.

Returns:

A dataset with more examples that are shorter.

Return type:

IQDataset

Examples:

If the dataset has 100 examples of length 1024 and the split factor is 2, the resulting dataset will have 200 examples of 512. No samples have been discarded.

If the example dataset has 100 examples of length 1024 and the example length is 100, the resulting dataset will have 1000 examples of length 100. The remaining 24 samples from each example have been discarded.

class ria_toolkit_oss.data.datasets.SpectDataset(source)[source]

Bases: RadioDataset, ABC

A SpectDataset is a RadioDataset tailored for machine learning tasks that involve processing radiofrequency (RF) signals represented as spectrograms. This class is integrated with vision frameworks, allowing you to leverage models and techniques from the field of computer vision for analyzing and processing radio signal spectrograms.

For machine learning tasks that involve processing on IQ samples, please use ria_toolkit_oss.data.datasets.IQDataset instead.

This is an abstract interface defining common properties and behaviour of IQDatasets. Therefore, this class should not be instantiated directly. Instead, it is subclassed to define custom interfaces for specific machine learning backends.

Parameters:

source (str or PathLike) – Path to the dataset source file. For more information on dataset source files and their format, see Intro to radio datasets.

property shape

Spectrogram datasets are M x C x H x W, where M is the number of examples, C is the number of image channels, H is the height of the spectrogram, and W is the width of the spectrogram.

Returns:

The shape of the dataset. The elements of the shape tuple give the lengths of the corresponding dataset dimensions.

Type:

tuple of ints

default_augmentations()[source]

Returns the list of default augmentations for spectrogram datasets.

Todo

This method is not yet implemented.

Returns:

A list of default augmentations.

Return type:

list[callable]

class ria_toolkit_oss.data.datasets.DatasetBuilder[source]

Bases: ABC

Abstract interface for radio dataset builders. These builder produce radio datasets for common and project datasets related to radio science.

This class should not be instantiated directly. Instead, subclass it to define specific builders for different datasets.

property name
Returns:

The name of the dataset.

Type:

str

property author
Returns:

The author of the dataset.

Type:

str

property url
Returns:

The URL where the dataset was accessed.

Type:

str

property sha256
Returns:

The SHA256 checksum, or None if not set.

Type:

str

property md5
Returns:

The MD5 checksum, or None if not set.

Type:

str

property version
Returns:

The version identifier of the dataset.

Type:

Version Identifier

property latest_version
Returns:

The version identifier of the latest available version of the dataset, or None if not set.

Type:

Version Identifier or None

property license
Returns:

The dataset license information.

Type:

DatasetLicense

property info
Returns:

Information about the dataset including the name, author, and version of the dataset.

Return type:

dict

abstract download_and_prepare()[source]

Download and prepare the dataset for use as an HDF5 source file.

Once an HDF5 source file has been prepared, the downloaded files are deleted.

abstract as_dataset(backend)[source]

A factory method to manage the creation of radio datasets.

Parameters:

backend (str) – Backend framework to use (“pytorch” or “tensorflow”).

Note: Depending on your installation, not all backends may be available.

Returns:

A new RadioDataset based on the signal representation and specified backend.

Type:

RadioDataset

ria_toolkit_oss.data.datasets.split(dataset, lengths)[source]

Split a radio dataset into non-overlapping new datasets of given lengths.

Recordings are long-form tapes, which can be obtained either from a software-defined radio (SDR) or generated synthetically. Then, radio datasets are curated from collections of recordings by segmenting these longer-form tapes into shorter units called slices.

For each slice in the dataset, the metadata should include the unique ID of the recording from which the example was cut (‘rec_id’). To avoid leakage, all examples with the same ‘rec_id’ are assigned only to one of the new datasets. This ensures, for example, that slices cut from the same recording do not appear in both the training and test datasets.

This restriction makes it challenging to generate datasets with the exact lengths specified. To get as close as possible, this method uses a greedy algorithm, which assigns the recordings with the most slices first, working down to those with the fewest. This may not always provide a perfect split, but it works well in most practical cases.

This function is deterministic, meaning it will always produce the same split. For a random split, see ria_toolkit_oss.data.datasets.random_split.

Parameters:

dataset (RadioDataset) – Dataset to be split.

Param:

lengths: Lengths or fractions of splits to be produced. If given a list of fractions, the list should sum up to 1. The lengths will be computed automatically as floor(frac * len(dataset)) for each fraction provided, and any remainders will be distributed in round-robin fashion.

Returns:

List of radio datasets. The number of returned datasets will correspond to the length of the provided ‘lengths’ list.

Return type:

list of RadioDataset

Examples:

>>> import random
>>> import string
>>> import numpy as np
>>> import pandas as pd
>>> from ria_toolkit_oss.data.datasets import split

First, let’s generate some random data:

>>> shape = (24, 1, 1024)  # 24 examples, each of length 1024
>>> real_part, imag_part = numpy.random.randint(0, 12, size=shape), numpy.random.randint(0, 79, size=shape)
>>> data = real_part + 1j * imag_part

Then, a list of recording IDs. Let’s pretend this data was cut from 4 separate recordings:

>>> rec_id_options = [''.join(random.choices(string.ascii_lowercase + string.digits, k=256)) for _ in range(4)]
>>> rec_id = [numpy.random.choice(rec_id_options) for _ in range(shape[0])]

Using this data and metadata, let’s initialize a dataset:

>>> metadata = pd.DataFrame(data={"rec_id": rec_id}).to_records(index=False)
>>> fid = os.path.join(os.getcwd(), "source_file.hdf5")
>>> ds = RadioDataset(source=fid)

Finally, let’s do an 80/20 train-test split:

>>> train_ds, test_ds = split(ds, lengths=[0.8, 0.2])
ria_toolkit_oss.data.datasets.random_split(dataset, lengths, generator=None)[source]

Randomly split a radio dataset into non-overlapping new datasets of given lengths.

Recordings are long-form tapes, which can be obtained either from a software-defined radio (SDR) or generated synthetically. Then, radio datasets are curated from collections of recordings by segmenting these longer-form tapes into shorter units called slices.

For each slice in the dataset, the metadata should include the unique recording ID (‘rec_id’) of the recording from which the example was cut. To avoid leakage, all examples with the same ‘rec_id’ are assigned only to one of the new datasets. This ensures, for example, that slices cut from the same recording do not appear in both the training and test datasets.

This restriction makes it unlikely that a random split will produce datasets with the exact lengths specified. If it is important to ensure the closest possible split, consider using ria_toolkit_oss.data.datasets.split instead.

Parameters:
  • dataset (RadioDataset) – Dataset to be split.

  • generator (NumPy Generator Object, optional.) – Random generator. Defaults to None.

Param:

lengths: Lengths or fractions of splits to be produced. If given a list of fractions, the list should sum up to 1. The lengths will be computed automatically as floor(frac * len(dataset)) for each fraction provided, and any remainders will be distributed in round-robin fashion.

Returns:

List of radio datasets. The number of returned datasets will correspond to the length of the provided ‘lengths’ list.

Return type:

list of RadioDataset

See Also:

ria_toolkit_oss.data.datasets.split: Usage is the same as for random_split().