Digital Microphone (DMIC) Acoustic Calibration & Array Tuning Guide#
Digital Microphone (DMIC) interfaces represent the primary audio capture ingress for modern Intel-based computing platforms, including Tiger Lake (cAVS 2.5), Arrow Lake (ACE 1.5), and Panther Lake (ACE 3.0). Unlike conventional analog microphone inputs that rely on external codecs, digital MEMS (Micro-Electro-Mechanical Systems) microphones integrate the acoustic transducer, preamplifier, and a 4th-order or 5th-order Sigma-Delta (\(\Sigma\Delta\)) modulator directly into a sub-millimeter silicon package. The transducers stream 1-bit oversampled Pulse Density Modulation (PDM) data directly into the DSP hardware.
The Sound Open Firmware (SOF) DMIC processing subsystem incorporates high-performance hardware decimation engines, programmable clock generation, dual-FIFO multirate dispatchers, and acoustic sensitivity/phase calibration filters. This guide provides an authoritative mathematical, architectural, and operational reference for configuring DMIC hardware decimators, calculating clock dividers and duty cycles, aligning microphone array phase and gain, compiling ACPI NHLT and ALSA Topology 2 binaries, and validating live transducer performance on target Hardware Under Test (DUT).
Transducer Physics & PDM Digital Ingress#
MEMS Digital Transducer Architecture#
Modern digital microphones convert ambient acoustic pressure fluctuations \(P(t)\) (measured in Pascals, where \(1\text{ Pa} = 94\text{ dBSPL}\)) into a high-rate 1-bit PDM pulse stream. The acoustic sensor consists of a flexible conductive diaphragm suspended over a rigid perforated backplate, forming a variable capacitor. Sound waves passing through the acoustic port deflect the diaphragm, modulating the capacitance.
An on-chip Application-Specific Integrated Circuit (ASIC) amplifies the capacitive charge and converts the continuous analog voltage into a 1-bit digital bitstream using an oversampled \(\Sigma\Delta\) modulator. The key acoustic metrology parameters governing digital microphones are summarized below:
Acoustic Sensitivity (\(S\)): The electrical signal level output by the microphone when subjected to a standard reference sound pressure level of \(1.0\text{ Pa}\) (\(94\text{ dBSPL}\)) at \(1\text{ kHz}\). For digital microphones, sensitivity is expressed in decibels relative to full scale (dBFS). A standard digital microphone has a nominal sensitivity of \(-26\text{ dBFS}\) (with typical production distributions ranging from \(-38\text{ dBFS}\) to \(-18\text{ dBFS}\)).
Acoustic Overload Point (AOP): The maximum sound pressure level at which the total harmonic distortion (THD) reaches \(10\%\) (or \(1\%\) depending on manufacturer rating). Standard mobile microphones offer an AOP of \(120\text{ dBSPL}\) to \(135\text{ dBSPL}\). Inputs exceeding the AOP cause severe clipping and \(\Sigma\Delta\) modulator saturation.
Signal-to-Noise Ratio (SNR): The ratio between the nominal sensitivity level (\(94\text{ dBSPL}\) at \(1\text{ kHz}\)) and the A-weighted acoustic noise floor of the microphone (\(V_{\text{noise}}\)), measured in \(\text{dBA}\):
\[\text{SNR} = 94\text{ dBSPL} - \text{EIN}\]Where \(\text{EIN}\) is the Equivalent Input Noise level in \(\text{dBSPL(A)}\). High-fidelity laptop arrays typically utilize microphones with an SNR between \(64\text{ dBA}\) and \(72\text{ dBA}\) (\(\text{EIN} \approx 22 - 30\text{ dBSPL}\)).
Double-Data-Rate (DDR) Stereo Multiplexing#
To minimize physical pin count and routing complexity across narrow laptop display hinges, digital microphones
use a shared two-wire interface consisting of a single clock line (PDM_CLK) and a single data line
(PDM_DATA). A pair of microphones (Mic A and Mic B) multiplexes onto this single data line using
Double-Data-Rate (DDR) timing governed by an external hardware SELECT pin:
Component |
SELECT Pin State |
Clock Driving Edge |
Bus Release Phase |
|---|---|---|---|
Mic A (L) |
Tied to Ground |
Rising Edge (Polarity 0) |
High-Z on Falling Edge |
Mic B (R) |
Tied to VDD |
Falling Edge (Polarity 1) |
High-Z on Rising Edge |
When PDM_CLK rises, Mic A latches its instantaneous 1-bit comparator state onto PDM_DATA while Mic B
maintains high-impedance (tri-state). When PDM_CLK falls, Mic A releases the bus into high-impedance,
and Mic B drives its 1-bit state onto PDM_DATA. The SOF hardware receiver samples both edges, de-interleaving
the stream into independent Left and Right channels.
Hardware Decimation Pipeline Architecture#
The DSP DMIC hardware controller ingests raw 1-bit PDM streams from up to 4 physical PDM controllers (supporting up to 8 microphone channels) and decimates them to linear 24-bit or 32-bit PCM audio samples. The complete signal chain is illustrated in Figure 256.
Figure 256 Digital Microphone (DMIC) Hardware Decimation Signal Processing Pipeline#
The signal processing chain consists of five sequential hardware stages:
Cascaded Integrator-Comb (CIC) 5th-Order Filter: High-ratio coarse decimation stage downsampling the overclocked 1-bit stream (\(f_{\text{pdm}}\)) to an intermediate rate (\(f_{\text{cic}}\)).
Arithmetic Shifter & Headroom Normalizer: Bit-alignment logic mapping the 26-bit CIC accumulator into the 22-bit input word of the FIR stage while preventing fixed-point overflow.
Finite Impulse Response (FIR) Multirate Filter: Precision shaping filter inverting the \(\text{sinc}^5\) passband droop of the CIC filter and downsampling to the target audio sample rate (\(f_s\)).
DC-Offset Compensation (DCCOMP): First-order high-pass Infinite Impulse Response (IIR) filter eliminating transducer DC bias and thermal drift.
Channel Gain Multipliers: 20-bit scaling registers balancing acoustic sensitivities across all elements of the microphone array.
Stage 1: 5th-Order Cascaded Integrator-Comb (CIC) Filter#
The primary decimation stage is implemented as a 5th-order Cascaded Integrator-Comb (CIC) filter. Because CIC filters require no multiplier units (utilizing only adders, subtractors, and delay registers), they operate directly at the multi-megahertz PDM clock rate with minimal power dissipation.
The discrete-time transfer function of an \(N\)-th order CIC filter with decimation factor \(M_{\text{cic}}\) is:
In Intel cAVS and ACE DSP architectures, the filter order is fixed at \(N = 5\) (5 cascaded integrator stages followed by 5 cascaded comb stages). The decimation factor \(M_{\text{cic}}\) is programmable between \(5 \le M_{\text{cic}} \le 31\).
The continuous-frequency magnitude response of the 5th-order CIC filter normalized to \(f_{\text{pdm}}\) is:
At zero frequency (\(f = 0\)), the DC power gain of the filter is:
Because \(M_{\text{cic}} \le 31\), the maximum theoretical DC gain is \(31^5 = 28,629,151\) (\(\approx 149.1\text{ dB}\)). The maximum bit growth through the five integration stages is:
For \(M_{\text{cic}} = 31\), \(B_{\text{growth}} = \lceil 5 \times 4.954 \rceil = 25\text{ bits}\). Including the
1-bit input sign, the internal accumulator requires 26 bits of precision, which matches the hardware width
constant DMIC_HW_BITS_CIC = 26.
Stage 2: Shifter Arithmetic & Headroom Normalization#
The FIR decimation engine expects signed 22-bit inputs (DMIC_HW_BITS_FIR_INPUT = 22). To bridge the
26-bit CIC accumulator to the 22-bit FIR input without clipping, an arithmetic right shifter scales the
CIC output. The required word length \(B_{\text{needed}}\) and the right shift offset are calculated by:
The shift value is programmed into the CIC_CONFIG register (bits 27:24) within the legal hardware range
\(-8 \le \text{cic\_shift} \le 4\). Because integer shifting attenuates signals by powers of 2 (\(2^{\text{cic\_shift}}\)),
the residual fractional gain headroom is transferred to the FIR coefficient scaling multiplier:
Stage 3: FIR Decimator & Sinc Droop Inversion#
While the CIC filter suppresses high-frequency quantization noise, it introduces a pronounced \(\text{sinc}^5\) gain droop across the audio passband:
At the edge of the passband (\(f = 20\text{ kHz}\) at \(f_s = 48\text{ kHz}\)), this droop attenuates high audio frequencies by up to \(-3.5\text{ dB}\) to \(-6.0\text{ dB}\), degrading vocal clarity and acoustic accuracy.
The second decimation stage employs a precision multirate FIR filter that performs two critical tasks:
Downsamples the audio stream by decimation factor \(M_{\text{fir}}\) (\(2 \le M_{\text{fir}} \le 15\)).
Equalizes the passband by implementing an exact inverse \(\text{sinc}^5\) frequency characteristic:
The FIR filters are configured with the following characteristics:
Passband Ripple: \(\le \pm 0.1\text{ dB}\) across \(0\text{ Hz}\) to \(0.4375 \times f_s\) (e.g. \(0 - 21\text{ kHz}\) at \(48\text{ kHz}\)).
Stopband Attenuation: \(\ge 90\text{ dB}\) to \(95\text{ dB}\) beginning at \(0.5100 \times f_s\), preventing alias reflection.
Coefficient Storage: Up to 250 filter taps (
DMIC_HW_FIR_LENGTH_MAX = 250) stored in dedicated SRAM as 20-bit signed integers (DMIC_HW_BITS_FIR_COEF = 20). Symmetric linear-phase filters exploit symmetry to store only \(\lceil N_{\text{taps}} / 2 \rceil\) unique coefficients.
Pipeline Hardware Cycle Constraints#
The FIR engine shares computational MAC units across active channels. For every audio output frame, the maximum number of FIR taps \(N_{\text{taps}}\) is bounded by the ratio between the DSP IO clock frequency (\(f_{\text{io}}\)) and the output sample rate (\(f_s\)):
The subtraction of 5 cycles represents the internal pipeline reload overhead (DMIC_FIR_PIPELINE_OVERHEAD = 5).
For standard clock configurations:
At \(f_{\text{io}} = 19.2\text{ MHz}\) and \(f_s = 48\text{ kHz}\): \(N_{\text{taps}} \le \min(250, \, 200 - 5) = 195\text{ taps}\).
At \(f_{\text{io}} = 38.4\text{ MHz}\) and \(f_s = 48\text{ kHz}\): \(N_{\text{taps}} \le \min(250, \, 400 - 5) = 250\text{ taps}\) (clamped at hardware maximum).
Stage 4: DC-Offset Compensation (DCCOMP)#
Digital MEMS microphones frequently exhibit intrinsic DC offsets resulting from diaphragm mechanical bias, \(\Sigma\Delta\) integrator leakage, and thermal drift. Uncompensated DC bias restricts downstream dynamic headroom, induces audible clicks and pops during stream starts, and corrupts time-domain acoustic feature extractors.
The hardware incorporates an independent first-order Infinite Impulse Response (IIR) DC-blocking high-pass filter on each channel:
The feedback pole \(\alpha = 1 - 2^{-k}\) determines the high-pass cutoff frequency \(f_c \approx 2^{-k} \cdot f_s / (2\pi)\).
The hardware provides 8 selectable time constants (DCCOMP_TC0 to DCCOMP_TC7):
Time Constant |
Bit Shift (\(k\)) |
Cutoff Frequency @ 48 kHz |
Settling Time (\(\tau\)) |
|---|---|---|---|
TC0 |
5 |
\(238.7\text{ Hz}\) |
\(0.67\text{ ms}\) (Fastest settling) |
TC1 |
6 |
\(119.4\text{ Hz}\) |
\(1.33\text{ ms}\) |
TC2 |
7 |
\(59.7\text{ Hz}\) |
\(2.67\text{ ms}\) |
TC3 |
8 |
\(29.8\text{ Hz}\) |
\(5.33\text{ ms}\) |
TC4 |
9 |
\(14.9\text{ Hz}\) |
\(10.67\text{ ms}\) |
TC5 |
10 |
\(7.5\text{ Hz}\) |
\(21.33\text{ ms}\) (Production default) |
TC6 |
11 |
\(3.7\text{ Hz}\) |
\(42.67\text{ ms}\) |
TC7 |
12 |
\(1.9\text{ Hz}\) |
\(85.33\text{ ms}\) (Infrasonic audio) |
Stage 5: Channel Output Gain Trimming#
Acoustic enclosures, cosmetic mesh grilles, and manufacturing tolerances introduce sensitivity variations
between microphone capsules. To present a balanced multichannel stream to downstream beamformers, the hardware
provides a dedicated 20-bit linear gain multiplier register for each channel: OUT_GAIN_LEFT_A,
OUT_GAIN_RIGHT_A, OUT_GAIN_LEFT_B, and OUT_GAIN_RIGHT_B.
The registers are encoded in unsigned \(Q1.19\) fixed-point format (where \(1.0\text{ (unity gain)} = 2^{19} = 524,288 = \text{0x080000}\)). For a target acoustic trim of \(\Delta G\text{ dB}\), the register value is:
Clock Generation & Dual-FIFO Architecture#
Primary Clock Division & Duty-Cycle Constraints#
The DSP clock generation block divides the platform main IO clock (\(f_{\text{io}} = 19.2\text{ MHz}\) on
older cAVS platforms or \(38.4\text{ MHz}\) on ACE platforms) to generate the physical PDM_CLK:
The 8-bit divider parameter is programmed into the MIC_CONTROL register as \(\text{PDM\_CLKDIV} = \text{clkdiv} - 2\).
The relationship between main IO clock, divider, and resulting PDM frequencies is illustrated in Figure 257.
Figure 257 DMIC Clock Generation and Dual-FIFO Mode Matching Architecture#
Odd dividers generate asymmetric clock high and low periods, altering the clock duty cycle:
MEMS microphone datasheets enforce strict duty cycle operational limits, typically \(40\% \le D \le 60\%\). If an odd divider produces a duty cycle outside this range (e.g. \(\text{clkdiv} = 3 \implies D_{\text{min}} = 33.3\%\)), the \(\Sigma\Delta\) modulator comparator timing fails, leading to noise floor rise or phase distortion. Furthermore, in cAVS 1.5 to 2.5 hardware, \(\text{clkdiv} \le 4\) is strictly prohibited by hardware timing paths.
Dual-FIFO Multirate Mode Matching#
A platform often requires two concurrent capture streams operating at different sample rates:
FIFO A (Communications / Recording): High-fidelity capture at \(f_{s,\text{A}} = 48\text{ kHz}\).
FIFO B (Voice Wake / Keyword Detection): Ultra-low-power processing at \(f_{s,\text{B}} = 16\text{ kHz}\).
Because both FIFOs receive data from the same physical microphones, they must share the exact same PDM clock frequency (\(f_{\text{pdm}}\)) and the exact same CIC decimation factor (\(M_{\text{cic}}\)). The multirate adaptation is achieved entirely within the FIR decimation stage by selecting different decimation factors \(M_{\text{fir,A}}\) and \(M_{\text{fir,B}}\):
Taking the standard ratio between \(48\text{ kHz}\) and \(16\text{ kHz}\) (\(3:1\)):
For example, on a platform with \(f_{\text{io}} = 38.4\text{ MHz}\):
Select \(\text{clkdiv} = 16 \implies f_{\text{pdm}} = 38.4\text{ MHz} / 16 = 2.40\text{ MHz}\) (Duty cycle = \(50.0\%\)).
Select \(M_{\text{cic}} = 25 \implies f_{\text{cic}} = 2.40\text{ MHz} / 25 = 96\text{ kHz}\).
For FIFO A (\(48\text{ kHz}\)): Select \(M_{\text{fir,A}} = 2 \implies 96\text{ kHz} / 2 = 48\text{ kHz}\). Total \(\text{OSR} = 25 \times 2 = 50\).
For FIFO B (\(16\text{ kHz}\)): Select \(M_{\text{fir,B}} = 6 \implies 96\text{ kHz} / 6 = 16\text{ kHz}\). Total \(\text{OSR} = 25 \times 6 = 150\).
Both streams run concurrently from a single physical PDM wire pair without clock conflict.
Microphone Array Acoustic Calibration & Phase Matching#
Array Geometry & Spatial Directivity#
Modern laptops and smart devices combine multiple digital microphones into spatial arrays to run Time-Domain Filter-and-Sum Beamformers (TDFB) or Real-Time Noise Reduction (RTNR) algorithms. The spatial geometry and propagation delays are illustrated in Figure 258.
Figure 258 Microphone Array Acoustic Geometry and Inter-Channel Phase Matching#
For two microphones spaced by distance \(d\), an acoustic plane wave arriving at incident angle \(\theta\) (where \(\theta = 0^\circ\) corresponds to broadside on-axis) experiences a physical propagation delay of:
Where \(c = 343\text{ m/s}\) is the speed of sound in air at \(20^\circ\text{C}\). The corresponding frequency-dependent phase shift is:
Impact of Acoustic & Transducer Mismatch#
Spatial beamformers create directional beams and steerable nulls by forming linear combinations of delayed microphone signals:
To place a deep null in the direction of ambient noise (\(\theta_{\text{null}}\)), the beamformer weights are designed so that \(W_1(f) X_1(f) + W_2(f) X_2(f) = 0\).
However, real-world hardware introduces two major sources of mismatch:
Magnitude Imbalance (\(\Delta G\)): Component manufacturing tolerances cause \(\pm 1.0\text{ dB}\) to \(\pm 1.5\text{ dB}\) sensitivity variations. Cosmetic acoustic mesh resistance and port hole dust seals introduce further attenuation differences.
Phase Skew (\(\Delta\phi\)): Acoustic port cavities act as acoustic low-pass Helmholtz resonators. Minor dimensional deviations in adhesive gasket thickness or acoustic port diameter shift the resonant frequency, introducing up to \(10^\circ\) to \(15^\circ\) of inter-channel phase error at \(4\text{ kHz}\) to \(8\text{ kHz}\).
As shown in Figure 258, a gain mismatch of just \(1.0\text{ dB}\) degrades spatial null depth from \(> 28\text{ dB}\) down to \(< 11\text{ dB}\), allowing ambient office noise, keyboard clicks, and echo to leak directly into the voice stream.
Acoustic Calibration Laboratory Setup#
To eliminate channel imbalance, microphone arrays must undergo acoustic calibration in a controlled metrology environment, illustrated in Figure 259.
Figure 259 Digital Microphone Acoustic Metrology and Sensitivity Calibration Rig#
The calibration test fixture requires:
Anechoic Test Box / Enclosure: Sound-isolated acoustic chamber providing \(\ge 40\text{ dB}\) ambient noise attenuation and lined with acoustic wedges to eliminate boundary reflections above \(200\text{ Hz}\).
Calibrated Reference Sound Source: Coaxial loudspeaker located at a fixed distance (\(d = 0.5\text{ m}\) or \(1.0\text{ m}\)) on the broadside axis (\(\theta = 0^\circ\)).
Class 1 Reference Microphone: Precision measurement microphone (e.g. Brüel & Kjær Type 4190 or GRAS 40AZ) calibrated using an acoustic calibrator to \(94.0\text{ dBSPL} \pm 0.1\text{ dB}\) at \(1\text{ kHz}\).
Audio Precision APx555 Analyzer: Precision audio generator driving the sound source and recording the reference microphone return signal.
Target Device Under Test (DUT): Connected via Ethernet or USB to record the uncalibrated multichannel DMIC PCM stream from SOF.
Sensitivity Calibration Procedure#
SPL Normalization: The Audio Precision analyzer plays a \(1\text{ kHz}\) sine wave and adjusts generator output until the reference microphone measures exactly \(94.0\text{ dBSPL}\) at the DUT position.
Raw Sensitivity Acquisition: Record a 5-second PCM capture from the DUT at \(48\text{ kHz}\) across all channels. Calculate the RMS digital level \(V_{\text{rms}, i}\) for each channel \(i\):
\[S_{\text{meas}, i} = 20 \log_{10}\left( \frac{V_{\text{rms}, i}}{V_{\text{FS}}} \right) \quad [\text{dBFS}]\]Gain Trim Calculation: Given a target nominal sensitivity \(S_{\text{target}}\) (typically \(-26.0\text{ dBFS}\)):
\[\Delta G_i = S_{\text{target}} - S_{\text{meas}, i} \quad [\text{dB}]\]\[\text{Trim}_{\text{linear}, i} = 10^{\frac{\Delta G_i}{20}}\]Register Programming:
Hardware \(Q1.19\) Register: \(\text{OUT\_GAIN}_i = \text{round}(\text{Trim}_{\text{linear}, i} \times 524,288)\).
IPC4 Copier \(Q10\) Parameter: \(\text{gain\_coeffs}[i] = \text{round}(\text{Trim}_{\text{linear}, i} \times 1024)\).
Post-calibration measurement must verify that channel sensitivity spread is within \(\le 0.05\text{ dB}\) across all elements.
Control Plane ABI & Topology 2 Integration#
ACPI NHLT (Non-HD-Audio Link Table) Configuration#
On Intel platforms, BIOS passes initial hardware decimation settings, clock dividers, and microphone array
geometry to the OS kernel via the ACPI NHLT (Non-HD-Audio Link Table). The table contains an endpoint
descriptor for the DMIC gateway, embedding the struct dmic_config_blob defined in src/include/ipc4/dmic.h:
/* Excerpt from src/include/ipc4/dmic.h */
struct dmic_config_blob {
uint32_t ts_group[4]; /* Time-slot channel mappings */
union dmic_global_cfg global_cfg;/* Clock-on delay & unmute fade settings */
uint32_t channel_ctrl_mask : 8; /* Active PDM channels */
uint32_t clock_source : 8; /* DSP IO clock selection (19.2M / 38.4M) */
uint32_t rsvd : 16;
struct dmic_channel_cfg channel_cfg[0];
uint32_t pdm_ctrl_mask; /* Bitmask of active PDM controllers (1..4) */
struct dmic_pdm_ctrl_cfg pdm_ctrl_cfg[0];
} __packed __aligned(4);
The nested struct dmic_pdm_ctrl_cfg holds the exact register images for each PDM controller:
cic_control: Soft reset, MIC A/B polarity, stereo mode.cic_config:COMB_COUNT(\(M_{\text{cic}} - 1\)) andCIC_SHIFT.mic_control:PDM_CLKDIV(\(\text{clkdiv} - 2\)) and clock edge selection.fir_config[2]: Length, decimation factor, DC offset, and channel gains for FIR A and B.fir_coeffs[0]: Array of 20-bit FIR coefficients (or packed 24-bit representations).
ALSA Topology 2 Declarations#
In Sound Open Firmware Topology 2, digital microphone DAIs are instantiated using Class.Dai."DMIC"
defined in tools/topology/topology2/include/dais/dmic.conf. A production topology declaration specifies:
# Production Topology 2 DMIC DAI instantiation
Object.Dai.DMIC."0" {
name "dmic01"
dai_index 0
direction "capture"
driver_version 1
io_clk 38400000
sample_rate 48000
clk_min 1000000
clk_max 4800000
duty_min 40
duty_max 60
num_pdm_active 2
fifo_word_length 32
unmute_ramp_time_ms 200
Object.Base.hw_config."0" {
id 0
}
Object.Base.pdm_config."0" {
ctrl_id 0
mic_a_enable 1
mic_b_enable 1
}
Object.Base.pdm_config."1" {
ctrl_id 1
mic_a_enable 1
mic_b_enable 1
}
}
Hardware Unmute Logarithmic Gain Ramp#
When digital microphones power on and clocking begins, capacitive charge stabilization in the MEMS capsule generates a low-frequency transient voltage thump. To eliminate audible pops, the SOF DMIC driver applies an automated logarithmic unmute gain ramp:
In src/include/sof/drivers/dmic.h:
LOGRAMP_START_DB= \(-90\text{ dB}\) (starting gain).Linear ramp equation: \(T_{\text{ramp}} = 200\text{ ms}\) at \(48\text{ kHz}\) and \(400\text{ ms}\) at \(16\text{ kHz}\).
Hardware unmute triggers: Unmute CIC at \(1\text{ ms}\) (
DMIC_UNMUTE_CIC = 1) and unmute FIR at \(2\text{ ms}\) (DMIC_UNMUTE_FIR = 2).
IPC4 Copier Runtime Gain Control#
Under IPC4, runtime acoustic trims can be injected without rebuilding the BIOS NHLT table. The driver
sends a DMA_CONTROL IPC containing the DMIC_SET_GAIN_COEFFICIENTS TLV (Type 2):
Field |
Byte Offset |
Type / Format |
Description |
|---|---|---|---|
Type |
0x00 |
uint32 (2) |
|
Length |
0x04 |
uint32 (8) |
Payload length in bytes (8 bytes) |
gain_coeffs[0] |
0x08 |
uint16 (Q10) |
Channel 0 Gain Trim (\(1.0 = 1024\)) |
gain_coeffs[1] |
0x0A |
uint16 (Q10) |
Channel 1 Gain Trim (\(1.0 = 1024\)) |
gain_coeffs[2] |
0x0C |
uint16 (Q10) |
Channel 2 Gain Trim (\(1.0 = 1024\)) |
gain_coeffs[3] |
0x0E |
uint16 (Q10) |
Channel 3 Gain Trim (\(1.0 = 1024\)) |
Standalone Python Calibration CLI Tool#
To streamline decimation parameter calculation, mode matching, and gain trim packaging, SOF provides the
standalone CLI utility sof_dmic_tool.py located in tools/tune/dmic/.
Searching Decimation Modes#
To evaluate all valid single or dual-rate decimation modes for a platform IO clock:
# Search matched modes for 48 kHz (comm) and 16 kHz (voice wake) on a 38.4 MHz platform
python3 tools/tune/dmic/sof_dmic_tool.py modes --ioclk 38.4e6 --rates 48000,16000
Example Tool Output:
==============================================================================
SOF DMIC Decimation Mode Search (IO Clock = 38.40 MHz)
==============================================================================
Dual-FIFO Matched Modes: FIFO A = 48000 Hz, FIFO B = 16000 Hz (Found 3 matches):
Idx clkdiv PDM Clock Duty M_cic M_fir_A M_fir_B CIC Shift
------------------------------------------------------------------------------
0 8 4.80 MHz 50.0% 25 4 12 3
1 10 3.84 MHz 50.0% 20 4 12 1
2 16 2.40 MHz 50.0% 25 2 6 3
==============================================================================
Calculating Sensitivity Trims & Building Binary Blobs#
Given laboratory sensitivity measurements across a 4-channel microphone array:
# Calibrate measured sensitivities to a target of -26.0 dBFS and output IPC4 binary blob
python3 tools/tune/dmic/sof_dmic_tool.py gain-trim \
--sens -25.2,-26.8,-24.9,-27.1 \
--target -26.0 \
--out dmic_gain_calibrated.bin
Example Tool Output:
==============================================================================
SOF Digital Microphone Acoustic Sensitivity Calibration
Target Sensitivity: -26.00 dBFS at 94 dBSPL (1 kHz)
==============================================================================
Ch Meas (dBFS) Trim (dB) Linear OUT_GAIN Reg (Q1.19) Copier (Q10)
------------------------------------------------------------------------------
0 -25.20 -0.80 0.9120 0x74BCC (478156 ) 0x03A6 (934)
1 -26.80 0.80 1.0965 0x8C596 (574870 ) 0x0463 (1123)
2 -24.90 -1.10 0.8810 0x70C63 (461923 ) 0x0386 (902)
3 -27.10 1.10 1.1350 0x91481 (595073 ) 0x048A (1162)
==============================================================================
Successfully generated binary IPC4 gain blob (16 bytes): dmic_gain_calibrated.bin
Production Calibration Recipes#
Recipe 1: Dual-Microphone Laptop Bezel (Broadside Array)#
Designed for standard clamshell and convertible laptops with two microphones spaced \(60\text{ mm}\) apart in the top display bezel.
Target Use-Case: High-definition video conferencing (Zoom, Teams) at \(48\text{ kHz}\) combined with background voice wake detection at \(16\text{ kHz}\).
Clock Architecture: \(f_{\text{io}} = 38.4\text{ MHz}\), \(\text{clkdiv} = 16 \implies f_{\text{pdm}} = 2.40\text{ MHz}\) (Duty cycle: \(50.0\%\)).
Filter Configuration: * \(M_{\text{cic}} = 25 \implies f_{\text{cic}} = 96\text{ kHz}\), \(\text{cic\_shift} = 3\). * FIFO A (\(48\text{ kHz}\)): \(M_{\text{fir,A}} = 2\) (Filter:
pdm_decim_int32_02, 63 taps). * FIFO B (\(16\text{ kHz}\)): \(M_{\text{fir,B}} = 6\) (Filter:pdm_decim_int32_06, 127 taps).DCCOMP: Mode
TC5(\(f_c = 7.5\text{ Hz}\)).Unmute Ramp: \(200\text{ ms}\) logarithmic ramp.
Recipe 2: Quad-Microphone Conference Tabletop Array (Circular)#
Designed for executive conference systems and smart hubs with 4 microphones arranged in a \(100\text{ mm}\) diameter circular geometry for \(360^\circ\) spatial speaker tracking.
Target Use-Case: 4-channel studio-quality capture with high acoustic overload ceiling (\(130\text{ dBSPL}\)).
Clock Architecture: \(f_{\text{io}} = 38.4\text{ MHz}\), \(\text{clkdiv} = 8 \implies f_{\text{pdm}} = 4.80\text{ MHz}\) (High performance mode).
Filter Configuration: * \(M_{\text{cic}} = 25 \implies f_{\text{cic}} = 192\text{ kHz}\), \(\text{cic\_shift} = 3\). * FIFO A (\(48\text{ kHz}\)): \(M_{\text{fir,A}} = 4\) (Filter:
pdm_decim_int32_04, 143 taps, Stopband: \(> 95\text{ dB}\)).DCCOMP: Mode
TC6(\(f_c = 3.7\text{ Hz}\)) for extended low-frequency vocal response.Sensitivity Alignment: Calibrated to \(-26.0\text{ dBFS} \pm 0.05\text{ dB}\) across all 4 channels.
Recipe 3: Ultra-Low-Power Edge Wake-on-Voice#
Designed for battery-constrained standby modes where the DSP monitors for keyword activation while drawing sub-milliwatt power.
Target Use-Case: Single or dual-mic keyword listening (\(16\text{ kHz}\)).
Clock Architecture: \(f_{\text{io}} = 19.2\text{ MHz}\), \(\text{clkdiv} = 25 \implies f_{\text{pdm}} = 768\text{ kHz}\) (Ultra-low-power mode, \(D = 48.0\%\)).
Filter Configuration: * \(M_{\text{cic}} = 16 \implies f_{\text{cic}} = 48\text{ kHz}\), \(\text{cic\_shift} = 0\). * FIFO B (\(16\text{ kHz}\)): \(M_{\text{fir}} = 3\) (Filter:
pdm_decim_int32_03, 45 taps).Power Dissipation: \(< 1.2\text{ mW}\) total digital subsystem power.
End-to-End Tuning Toolchain Workflow#
The complete end-to-end DMIC engineering workflow is illustrated in Figure 260.
Figure 260 End-to-End DMIC Tuning and Calibration Toolchain Workflow#
The workflow encompasses 5 coordinated stages:
Hardware Specification: Reviewing microphone datasheet limits (PDM clock min/max, duty cycle tolerances, sensitivity, AOP) and physical acoustic port enclosure geometry.
Filter Tuning & Mode Selection: Running
sof_dmic_tool.pyor Octavedmic_init.mto select valid integer decimation tuples and generate droop-compensating FIR filter taps.Acoustic Calibration: Measuring the DUT array in an anechoic box with an Audio Precision APx555, deriving channel sensitivity deltas, and calculating \(Q1.19\) and \(Q10\) gain trim coefficients.
Blob Packaging & Compilation: Populating ACPI NHLT descriptors and ALSA Topology 2 configuration files, then compiling target binary artifacts (
.tplgandnhlt-*.bin).Target Deployment & Sign-Off: Deploying binaries to the DUT, verifying live streams via
arecordandsof-ctl, and confirming that THD+N, frequency flatness, and beamformer directivity satisfy requirements.
Interactive Live Injection & Diagnostics Matrix#
Runtime Gain Verification via sof-ctl#
To inspect or inject digital microphone gain trims on a live DUT over SSH:
# Step 1: Query active mixer controls on the DMIC capture card
ssh root@<dut> "amixer -c 0 scontrols | grep -i dmic"
# Step 2: Set capture volume via ALSA mixer (in decibels)
ssh root@<dut> "amixer -c 0 sset 'DMIC01 Capture Volume' 20dB"
# Step 3: Inject binary gain calibration blob into active IPC4 copier component
# Widget ID 12 corresponds to the DMIC ingress copier
scp dmic_gain_calibrated.bin root@<dut>:/tmp/dmic_gain.bin
ssh root@<dut> "sof-ctl -D hw:0 -w 12 -s /tmp/dmic_gain.bin"
# Step 4: Record a 10-second multi-channel test capture to verify audio integrity
ssh root@<dut> "arecord -D hw:0,1 -f S32_LE -c 4 -r 48000 -d 10 /tmp/dmic_test.wav"
Diagnostic Troubleshooting Matrix#
Symptom |
Root Cause |
Diagnostic Command |
Remediation Action |
|---|---|---|---|
Audible Thump/Click on Stream Start |
Capsule DC bias during clock power-up ramp |
Inspect kernel dmesg: |
Increase |
Digital Clipping / Hard Saturation at Moderate SPL Levels |
CIC shifter underflow or gain multiplier overflow |
Analyze recorded WAV: peak at \(0\text{ dBFS}\) with flat tops |
Re-evaluate |
Severe Noise Floor Rise / Modulator Hash |
Non-compliant clock duty cycle (\(< 40\%\)) from odd |
Measure |
Avoid odd dividers that yield duty cycles outside \([40\%, 60\%]\); select higher \(f_{\text{io}}\) clock. |
Degraded Beamformer Null Depth (< 15 dB) |
Channel gain spread \(> 0.5\text{ dB}\) or acoustic port leakage |
Run APx555 sensitivity sweep or compare RMS power across recorded channels |
Re-run acoustic calibration in anechoic box; inject precise gain trims via |
180° Inverted Channel Polarity |
Inverted clock edge selection in hardware |
Inspect waveform polarity on dual-channel impulse stimulus |
Toggle |
FIFO Buffer Overrun / DSP Panic |
FIR tap count exceeds allowed MAC cycles per frame interval |
Check SOF trace log: |
Reduce FIR tap count (\(N_{\text{taps}} \le \lfloor f_{\text{io}} / (2 f_s) \rfloor - 5\)) or increase platform IO clock speed. |