Exploratory acoustic feature extraction, quality control, within-subject reference estimation, and longitudinal analysis for investigational voice research.
Version 1.1.0 · Last updated September 2026
Cogitrac Voice is an investigational research platform designed to extract acoustic and spectral measurements from voice recordings for longitudinal research. The platform supports exploratory investigation of whether voice-derived variables may be associated with within-subject changes or prospectively defined study outcomes.
Investigational research platform. The extracted measurements are candidate research variables. They have not been established as validated biomarkers of epilepsy, seizures, neurological disease, or any other clinical condition. Individual measurements should not be interpreted as diagnostic, prognostic, treatment-related, or clinically actionable results.
Feature extraction is intended to preserve potentially informative acoustic measurements for subsequent research. The fact that a feature is extracted or stored does not establish that it has clinical significance, predictive value, or a causal relationship with a study outcome.
| Component | Status | Interpretation |
|---|---|---|
| Mean F0 | Extracted | Candidate acoustic variable for research analysis. |
| F0 variability (std) | Extracted | Standard deviation of eligible detected voiced F0 values within the recording; candidate within-recording variability measure. |
| Jitter / shimmer | Conditional | Exploratory voice-quality measurements requiring sufficient valid periodic intervals after pitch and pulse detection. The current implementation additionally requires at least three detected pulses as a computational eligibility gate; this criterion does not establish measurement reliability. |
| Formants | Extracted | Candidate vocal-tract resonance measures (F1, F2, F3); estimates are subject to task, signal-quality, speaker, and algorithmic limitations. |
| MFCCs | Extracted | MFCC 1–13 (excluding coefficient 0) are extracted from short-time frames using a Mel-frequency filterbank with prespecified frame length, hop length, and windowing function. |
| Custom derived features | Derived | Exploratory variables requiring independent evaluation and validation. |
| Seizure detection | Not validated | No clinical seizure-detection claim is established by this methodology. |
| Feature-level analysis framework | Implemented | Supports exploratory within-subject analysis of feature behavior across recordings. |
| Clinical biomarker threshold | Not established | Requires appropriately designed development and independent validation studies. |
Voice measurements can be affected by recording equipment, device processing, microphone position, background noise, room acoustics, speaking task, language, hydration, fatigue, illness, medication, emotional state, and ordinary biological variation.
Consequently, longitudinal comparisons should be interpreted in the context of recording conditions and relevant protocol information.
Research consideration: A change in an acoustic measurement may reflect recording conditions, ordinary biological variation, task differences, or another confounding factor rather than a change in the study outcome of interest.
Longitudinal interpretation: Observed within-subject changes should not be interpreted as evidence of an association with a clinical or study outcome unless that association is evaluated using a prespecified statistical analysis and appropriate controls for relevant confounding and repeated measurements.
Recordings undergo basic quality checks before acoustic feature extraction. The current extraction implementation includes an overall signal-level check intended to identify recordings that are effectively silent or extremely low in amplitude.
np.nan and should be excluded or otherwise handled
according to the downstream analysis plan.
Recording-level vs. feature-level missingness: A recording may pass recording-level quality control while individual features remain undefined because feature-specific eligibility criteria are not satisfied. Feature-level missingness should be retained and handled according to the prespecified statistical analysis plan rather than automatically imputing values.
1e-5 in the current implementation, calculated after
amplitude normalization.
Implementation limitation: For the current study protocol, recordings shorter than 10 seconds are excluded prior to feature extraction. This is a protocol-defined operational criterion rather than a validated minimum duration for reliable estimation of the individual acoustic features. It does not by itself establish measurement reliability for any individual feature. Duration requirements may vary across study protocols and should be specified in the applicable study documentation.
Feature extraction uses Parselmouth, a Python interface to Praat, together with Librosa for complementary digital signal-processing features.
The current pipeline extracts a broad set of acoustic and spectral variables. These measurements are retained as candidate variables for exploratory longitudinal analysis, reproducibility studies, feature selection, and future research.
Interpretation: Extraction does not imply validation. A feature can be scientifically useful for hypothesis generation without having established clinical, diagnostic, prognostic, or neurological significance.
| Component | Configuration | Notes |
|---|---|---|
| Frame length | ~46 ms | Configured for spectral feature extraction |
| Hop length | ~12 ms | 75% frame overlap |
| Windowing | Hann | Applied prior to spectral analysis |
| MFCC coefficients | 1–13 | Coefficient 0 (energy) excluded |
| Formant tracking | Burg algorithm | Maximum of 4 formants evaluated; F1–F3 retained for the current feature set |
Configuration note: Exact parameter values (frame length, hop length, formant ceilings, and advanced tuning parameters) are maintained in the version-controlled production environment. Qualified researchers may request detailed specifications under appropriate confidentiality agreements.
Mean fundamental frequency across detected voiced pitch frames.
Parselmouth · ExploratoryStandard deviation of eligible detected voiced F0 values within the recording.
Parselmouth · ExploratoryAn exploratory measure of similarity between F0 trajectories in the first and second portions of the recording. Missing F0 values are excluded, and the two halves are trimmed to equal length. A prespecified minimum number of valid F0 frames is required, with a prespecified minimum in each half. If insufficient frames are available, the measure is undefined. This measure should be interpreted as a within-recording descriptive statistic rather than a validated measure of vocal consistency.
Derived feature · ExploratoryLocal cycle-to-cycle variation in detected period duration, calculated using the configured Praat local-jitter procedure.
Parselmouth / Praat · ExploratoryLocal cycle-to-cycle variation in detected period amplitude, calculated from the Sound and PointProcess representations.
Parselmouth / Praat · ExploratoryMean harmonicity-derived HNR across eligible analyzed time intervals.
Parselmouth / Praat · ExploratoryMeasurement caveat: Jitter and shimmer are established acoustic voice-quality measurements, but their reliability depends on recording characteristics, pitch tracking, voice type, task, and analysis parameters. They are not necessarily appropriate for all types of speech recordings, and their interpretation depends substantially on the type of vocal material and the conditions under which periodic cycles are detected. Accordingly, jitter and shimmer should not be interpreted as universally comparable across recording tasks or protocols without protocol-specific reliability assessment.
First-formant estimate obtained at the midpoint of the analyzed recording when a valid formant estimate is available.
Parselmouth · ExploratorySecond-formant estimate obtained at the midpoint of the analyzed recording when a valid formant estimate is available.
Parselmouth · ExploratoryThird-formant estimate obtained at the midpoint of the analyzed recording when a valid formant estimate is available.
Parselmouth · ExploratoryAn exploratory measure of within-recording F1 and F2 variability, calculated from formant estimates sampled across the recording duration. This is a descriptive acoustic statistic and should not be interpreted as a validated measure of articulatory instability.
Derived feature · ExploratoryFormant eligibility: Formant estimates are algorithm-dependent and derived from the full recording. Estimates are flagged as invalid when a valid formant estimate is unavailable under the configured analysis criteria. The exact formant-tracking parameter values are version-controlled and should be reported with analyses using these measurements.
Mean root-mean-square energy across short-time frames, calculated on the normalized waveform.
Librosa · Extracted / retainedStandard deviation of RMS energy across short-time frames, providing a measure of energy variability within the recording.
Derived feature · ExploratoryMean spectral centroid across short-time frames.
Librosa · Extracted / retainedMean spectral bandwidth across short-time frames.
Librosa · Extracted / retainedMean spectral rolloff frequency across short-time frames.
Librosa · Extracted / retainedMean zero-crossing rate across short-time frames.
Librosa · Extracted / retainedSimilarity measure between spectral magnitudes in the low-frequency and mid-frequency bands, derived from the frequency spectrum.
Derived feature · ExploratoryMel-frequency cepstral coefficients 1 through 13 (excluding coefficient 0) are extracted from short-time frames. Each coefficient is summarized at the recording level by its mean across eligible frames. Individual MFCC coefficients are retained as candidate variables for exploratory research and should not be interpreted as validated biomarkers.
Librosa · Extracted / retainedFeature status: The extraction layer is intentionally broader than any individual research hypothesis. Features may be retained for later analysis, feature selection, replication, and hypothesis generation. Inclusion in the extraction pipeline does not establish clinical significance, predictive value, or association with epilepsy, seizures, or any neurological outcome.
Where longitudinal analysis is performed, measurements may be compared with prior recordings from the same participant. The purpose of a within-subject reference is to characterize observed individual variation rather than assume that a population-level reference value is clinically meaningful for every participant.
Study-specific reference and event windows should be defined prospectively in the research protocol. For example, a study may designate recordings collected within a specified period before an independently documented event as event-proximal observations while using separate recordings for baseline estimation.
Important: Reference-window definitions are study-design decisions, not validated clinical rules. Where feasible, they should be specified before outcome analysis to reduce analytical flexibility and potential bias.
Baseline sample size should be determined prospectively based on the expected within-subject variability, desired precision, statistical model, number and frequency of measurements, and study objectives.
A within-subject reference may be estimated from prespecified baseline recordings meeting quality-control criteria. Reference estimates should be defined before outcome analysis and should not be selected retrospectively based on observed outcomes.
The reference distribution is characterized by its mean (μref) and standard deviation (σref) calculated from eligible baseline recordings. Recordings that fail quality control, contain excessive noise, or are otherwise non-representative of the participant's typical voice are excluded from reference estimation.
Reference independence: Where applicable, baseline/reference recordings should be selected independently of the outcome under investigation and according to prospectively defined temporal and clinical criteria. Recordings potentially influenced by the outcome or its immediate precursors should not be included in the reference set unless explicitly permitted by the study protocol.
If insufficient baseline recordings are available, the reference estimate may be based on available recordings, and the limitations should be documented in the study protocol.
Feature extraction and statistical analysis are separate stages of the research pipeline. The extraction layer produces acoustic measurements; downstream analyses evaluate those measurements using an exploratory statistical framework designed to support hypothesis generation across longitudinal voice data.
Framework overview: The analysis framework supports exploratory evaluation of individual observations relative to a within-subject reference, assessment of group-level feature behavior, and evaluation of whether combinations of feature changes recur across multiple observed events. Each stage produces candidate research variables, not clinical conclusions.
For exploratory analyses using within-subject standardization, individual measurements may be compared against the participant's own reference distribution. This produces a standardized deviation score describing how far a given measurement sits from the participant's typical range, expressed in units of within-subject variability.
Standardization caveat: When the participant's within-subject variability is zero or below a prespecified minimum threshold, a standardized score is considered undefined rather than assigning an arbitrarily large value. This avoids over-interpretation of scores based on near-zero reference variability.
Reference recomputation: Within-subject reference statistics are recomputed each time an analysis is run, using all eligible reference recordings available at that time. As additional recordings accumulate, the estimated reference distribution may change, and consequently the standardized score of any previously recorded measurement may also change on subsequent analyses. Longitudinal interpretation should account for this behavior.
A standardized score describes deviation from the selected reference distribution. It is not, by itself, a biomarker, diagnostic threshold, event probability, or measure of disease severity.
Downstream analyses may compare measurements from two prespecified groups of recordings, such as event-proximal recordings collected within a prespecified time window before an independently documented event, and reference recordings collected outside a prespecified exclusion window surrounding each event.
Comparisons are performed using standard non-parametric methods, with multiple-comparison correction applied across the feature set at a prespecified significance level. The specific statistical procedures, thresholds, and parameters used in the current implementation are maintained in the version-controlled production environment and are available to qualified researchers under appropriate confidentiality agreements.
Interpretation: Statistical assessment produces candidate research observations. A candidate observation indicates a systematic difference under the current data and does not by itself establish predictive value, causality, or clinical utility.
Longitudinal analyses should account, as appropriate, for repeated measurements within participants, temporal dependence, reference drift, missing observations, multiple comparisons, recording-condition effects, task differences, and relevant participant-level covariates.
Statistical thresholds and model specifications intended for confirmatory analyses should be defined prospectively where feasible. Exploratory findings should be clearly distinguished from confirmatory results.
Extracted features may be evaluated as candidate research variables based on statistical behavior, measurement reliability, reproducibility, association with prospectively defined study outcomes, temporal relationships, and robustness to relevant confounding factors.
Exploratory analyses may identify features warranting further investigation. An association observed in an exploratory dataset does not establish that a feature is a validated biomarker.
Important: Candidate-feature identification is an exploratory research step. Statistical association or predictive performance in an exploratory dataset does not establish clinical validity, clinical utility, or biomarker status. Independent validation is required before making such conclusions.
Features, thresholds, preprocessing decisions, and model specifications selected using exploratory or training data should not be evaluated on the same observations used to make those selections.
Where predictive modelling is performed, participant-level separation should generally be used so that recordings from the same participant do not appear in both development and validation datasets. This reduces the risk that measured performance reflects participant-specific memorization or within-participant leakage rather than generalization to new individuals.
Establishing a clinical biomarker or predictive model would require appropriately designed development and validation studies using prospectively defined outcomes, prespecified analysis procedures, participant-level separation of development and validation data, assessment of measurement reliability and reproducibility, evaluation of relevant confounding factors, and independent validation in an appropriate target population.
Important: Performance observed in exploratory or development data would not by itself establish clinical validity or clinical utility. Independent validation in appropriately separated data is required before any clinical or diagnostic conclusions can be drawn.
This methodology describes the processing pipeline associated with Cogitrac Voice version 1.1.0. Changes to preprocessing parameters, feature definitions, recording eligibility criteria, reference-window rules, or statistical thresholds may produce results that are not directly comparable across versions.
The production implementation and its version-controlled pipeline specifications are the definitive references for exact processing behavior.
The detailed audio preprocessing and feature-extraction pipeline, including parameter values, exclusion criteria, and quality-control thresholds, is maintained in the version-controlled production environment.
Key specifications include:
These specifications are documented internally and are available to qualified researchers under applicable confidentiality agreements.
Complete technical details are available to qualified researchers and collaborators where appropriate.
Contact Research TeamReference scope: These publications provide background on acoustic voice measurement, software, and signal processing. They do not constitute validation of the Cogitrac Voice feature set for epilepsy, seizures, or other neurological outcomes.