Skip to main content

PHOENIX USER MANUAL

Complete guide to using Phoenix

Introduction

Welcome to Phoenix — a hybrid virtual synthesizer created by BT and Spitfire Audio for composers, sound designers, and producers who want to go beyond a single mono oscillator per note. This guide walks through every part of the instrument in plain language, so whether you're opening your first preset or building a patch from scratch, you'll know what each section does and why it's useful.

Use the index below to jump straight to the section you need.

Contents


System Requirements

The instrument uses a unified hybrid synthesis architecture in which up to six synthesis engines can operate simultaneously: spectral wavetable synthesis, FM synthesis, sample playback, 2-operator FM transient synthesis, subharmonic generation and a complex noise oscillator.

The instrument also supports up to 16-voice full MPE.

CPU requirements are therefore highly dependent on the sound being played. A relatively simple preset may consume substantially less processing power than a sound using all six engines, extensive modulation, effects and high polyphony.

macOS — Minimum

Apple Silicon

  • macOS 15 or later

  • Apple M1 or later

  • 16 GB RAM

Intel

  • macOS 15 or later

  • 2019 or later Intel Mac

  • 8-core Intel Core i9 or better

  • 16 GB RAM

Storage and Host

  • SSD required

  • Approximately 80 GB free disk space for installation and library

  • 64-bit AU, VST3 or AAX compatible host

macOS — Recommended

  • Apple M2 Pro, M3 or newer

  • 32 GB RAM

  • Fast internal SSD or high-performance external SSD

  • 80 GB or more available disk space

  • Current supported version of macOS

  • 512-sample audio buffer for CPU-intensive operation

Apple Silicon is strongly recommended for users intending to use complex presets, extensive polyphony, 16-voice MPE or multiple simultaneous instances.

Intel Mac Performance

Intel Macs remain supported, but the instrument's hybrid synthesis architecture places substantial demands on real-time CPU performance.

The practical minimum supported Intel configuration is:

2019 or later Intel Mac with an 8-core Intel Core i9 processor and 16 GB RAM.

Older 6-core Intel systems, including Intel Core i7 configurations, may not provide sufficient real-time processing performance for demanding presets and are not included in the minimum supported specification.

Users operating supported Intel systems should expect to use larger audio buffer sizes with CPU-intensive presets.

A 512-sample buffer is recommended as the starting point for demanding sounds on Intel Mac systems.

Even on supported Intel systems, particularly demanding configurations may require reduced polyphony, fewer simultaneous instances or rendering/freezing of instrument tracks.

Meeting the minimum system specification does not guarantee that every possible synthesis configuration can be played at maximum polyphony and minimum audio-buffer latency.

Windows — Minimum

  • Windows 10 64-bit, version 22H2 or later

  • Intel Core i7 10th Generation or newer

or

  • AMD Ryzen 5 5000 series or newer

  • 16 GB RAM

  • SSD required

  • Approximately 50 GB free disk space for installation and library

  • 64-bit VST3 or AAX compatible host

Windows — Recommended

  • Windows 11 64-bit

  • Intel Core i7 12th Generation or newer

or

  • AMD Ryzen 7 5000 series or newer

  • 32 GB RAM

  • Fast NVMe SSD

  • 80 GB or more available disk space

  • 512-sample audio buffer for CPU-intensive operation

Supported Audio Configuration

Sample rate: 48 kHz

Plug-in formats

macOS:

  • Audio Units (AU)

  • VST3

  • AAX

Windows:

  • VST3

  • AAX

64-bit hosts are required.

Audio Buffer Size

Audio buffer size has a significant effect on the amount of processing time available to the instrument.

For general playback on a modern high-performance system, buffer sizes below 512 samples may operate successfully. The achievable setting depends on processor performance, DAW workload, audio interface and driver performance, preset complexity and polyphony.

For demanding cinematic presets, high-polyphony passages, extensive MPE performance or operation on systems close to the minimum specification, we recommend:

48 kHz / 512 samples

If CPU overloads, clicks, pops or interrupted playback occur:

  1. Increase the audio buffer to 512 samples.

  2. Reduce the preset voices where appropriate.

  3. Reduce the number of simultaneously active notes and long release tails.

  4. Freeze, bounce or render other CPU-intensive tracks.

  5. Close unnecessary applications and background processes.

  6. Ensure that the library is installed on a fast SSD.

Lower-Latency Recording

For live recording or performance where lower latency is important, a smaller buffer such as 128 or 256 samples may be used if the computer provides sufficient processing headroom.

Complex sounds that operate reliably at 512 samples may exceed the available real-time processing capacity at 128 or 64 samples.

For large arrangements, sound-design sessions and final playback, increasing the buffer to 512 samples is recommended where low monitoring latency is not required.

Understanding CPU Performance

Real-time synthesizer performance cannot be determined from total processor core count or processor clock speed alone.

DAWs distribute plug-in processing according to their own scheduling architecture, and individual real-time processing chains may be constrained by the performance available to a limited number of CPU cores.

Processor architecture, generation and single-core performance should therefore be considered alongside total core count and clock speed.

This is particularly relevant when comparing older Intel processors with Apple Silicon and newer-generation Windows processors.

MPE and Polyphony

The instrument supports up to 16-voice full MPE.

MPE can significantly increase processing requirements because individual voices may receive independent pitch, pressure, timbre and other modulation data. Complex modulation combined with multiple synthesis engines can therefore produce substantially greater CPU demand than conventional single-channel MIDI playback.

For demanding 16-voice MPE performances, we recommend:

  • A system meeting or exceeding the recommended CPU specification

  • 32 GB RAM

  • 512-sample audio buffer where low latency is not essential

  • Apple Silicon on macOS

Storage

The complete library requires more than 40 GB of storage.

An SSD is required for reliable operation. A fast internal SSD or NVMe drive is recommended.

Allow at least 80 GB of available disk space for the installed product, and preferably 100 GB or more to provide sufficient space for installation, updates and library management.

Mechanical hard drives are not recommended for the main library.

Performance Expectations

The minimum system requirements define a practical baseline for running the instrument. They should not be interpreted as guaranteeing maximum performance under every possible synthesis configuration.

CPU demand can increase considerably when combining:

  • Multiple simultaneous synthesis engines

  • All six synthesis engines simultaneously

  • High polyphony

  • 16-voice MPE

  • Complex modulation

  • Long release times

  • CPU-intensive effects

  • Multiple plug-in instances

For professional cinematic composition and sound-design sessions involving multiple instances, we recommend exceeding the minimum specification and using the recommended system configuration wherever possible.


What Makes Phoenix Different

Phoenix combines six sound-generating engines ("oscillators") inside every single voice:

  • A stereo wavetable engine

  • A 48-voice FM engine

  • A deep multi-sampled instrument engine, built on Spitfire Audio's recorded sample libraries

  • A sampled-FM transient (attack) engine

  • A sub-oscillator

  • A noise oscillator

All six share one modulation system and one signal path, so any of them can be shaped, blended, and modulated together. On top of that sits a large collection of built-in effects — 37 in total — each with its own dedicated screen rather than a generic set of knobs.

The idea that ties the whole instrument together is simple: everything is genuinely stereo, not just spread out afterwards. The next section explains what that means in practice.


Two Synthesizers in One: The Stereo Engine

Most synthesizers generate a single mono waveform and then create the impression of width afterwards — by detuning and panning multiple copies of a voice, adding chorus, or spreading amplitude across the field. The oscillator itself only ever produces one signal; everything stereo about the sound is added on top.

Phoenix works differently. Its wavetable oscillator runs two complete, independent synthesis chains — one per channel — all the way from the raw waveform to the final output. In practical terms, that means the left and right side of a single note can:

  • Sit at different positions in the wavetable

  • Carry different harmonic content

  • Use different phase-warp settings

  • Drift apart in pitch over time

Because of this, several core parameters — wavetable position, both phase-warp amounts, the spectral warp amount, transpose, and fine tune — can each be modulated independently per channel. Route a modulator to wavetable position, for example, and the two sides can scan the table at different speeds, landing on different harmonic content, rather than moving together.

Four dedicated tools exist purely to push the two channels apart on purpose:

  • Stereo remap presets — a number of the shaping curves have entirely independent left and right versions

  • Stagger — offsets the left and right read-heads in opposite directions

  • Mirror — rotates the two channels' phase apart, in opposite directions

  • Flip Pan — crossfades each channel toward the other channel's setting, for a controlled stereo inversion


The Oscillators

Wavetable Oscillator

The wavetable oscillator reshapes recorded or synthesized waveforms in real time, giving you a huge range of tones from a small starting table. It supports:

  • Up to 256 frames and 2,048 samples per frame

  • 1,024 usable harmonics, each individually accessible (see The Spectral Editor)

  • Extremely fine pitch resolution, well beyond what the ear can perceive

To keep everything clean at extreme settings, Phoenix runs four anti-aliasing systems at once in the background — technical processes that prevent the harsh, digital-sounding artifacts that aggressive waveform reshaping can otherwise cause. In everyday terms: you can push the oscillator hard without it turning into unwanted noise.

Phase Warping

Phase warping reshapes how the wavetable is read, rather than the table itself. It's the underlying idea behind classic techniques like hard sync, pulse-width modulation, and wavefolding — Phoenix treats all of them as one large, explorable design space.

  • 86 phase-warp modes, grouped into families: Basic, Remap, Sync, Distortion, Filter, PWM, and Advanced Warp

  • Two warp slots, in series — the order matters, so running a fold into a sync produces a different result than running a sync into a fold. With 74 of those modes available in each slot, there are well over five thousand possible combinations before you even touch a depth control

  • Both depth controls can be modulated independently per channel, so, for example, the left ear can be folding while the right ear is syncing

Sitting inside this system is a dedicated remap engine: 28 curve presets (each a detailed 129-point shape), four live controls that reshape any of them (Stagger, Mirror, Flip Pan, Curve Warp), and six additional character controls. Every curve can be hand-edited. The effect works like a resonator or formant filter built entirely from phase — shaping tone without adding a filter to the signal path.

Automatic Level Balancing

Reshaping a waveform aggressively normally causes big, unpredictable jumps in loudness — the usual fix is to add a compressor afterwards, which flattens out the very transient character that made the sound interesting in the first place.

Phoenix avoids this by measuring instead of compressing. Every warp mode, at every setting, was measured offline in advance (millions of individual data points), and that data is built into the instrument so it can automatically apply the exact right amount of gain compensation in real time. The practical result: you can sweep a warp control from subtle to extreme, stack two warps together, and pile on unison voices — and the volume stays consistent throughout, with no compressor in the way.

The Spectral Editor

The spectral editor works on a sound the way a prism works on light: it splits a note into its individual harmonics (partials) and gives you access to every one of them individually — decay time, drift, and position in the harmonic series.

  • Evocative new spectral modes ship as starting points — from struck glass and metal to ghostly, evolving textures — and every one is fully open and editable

  • A ratio family of controls bends the harmonic spacing itself: snap harmonics to the nearest whole number, nearest power of two, odd harmonics only, a Fibonacci sequence, prime numbers, or the golden ratio, among others

  • A set of physical templates re-tunes the harmonics to match real acoustic objects — a stretched piano tuning, a tubular bell, a marimba, a drumhead, and several just-intonation and equal-tempered tuning systems

  • A level family of controls (tilt, hype, comb, and more) reshapes the overall spectral balance

There's also a second, faster spectral stage with 9 additional spectral warp modes (such as Blur, Rounder, and Harmonic Smear) that can be modulated independently per channel — layering yet another type of harmonic movement on top of the main spectral engine.

The Wavetable Editor & Resynthesis

The wavetable editor gives you direct, hands-on control of the raw waveform:

  • Draw individual wave cycles by hand using dedicated brushes

  • Edit all 1,024 partials individually, in both amplitude and phase

  • Morph through up to 256 frames so a sound has a trajectory over time, not just a static shape

  • Generate frames from formulas, and sort, blur, or phase-align across an entire table at once

The standout feature here is resynthesis. Drag in a recording of almost any struck object — a marimba, a wine glass, a piece of metal — and Phoenix analyzes it, measuring the frequency, level, decay time, and stereo position of every partial in the recording. It then rebuilds that object as a fully playable instrument you can perform across the keyboard, with every measured partial available as a modulation target. You can bend the result gently (a "marimba made of glass") or push it much further into something that never existed acoustically, while it still behaves like a real struck object because the underlying model is physical, not a static sample.

FM Oscillator

Phoenix includes a dedicated six-operator FM engine (rather than treating FM as just another modulation routing option), and either synthesis slot can host one — so a single voice can run two independent six-operator FM oscillators at once.

  • Six operators, 34 algorithms — including a full set of algorithms compatible with the classic FM algorithm layouts many musicians already know, plus two original Phoenix algorithms (see appendix for a brand reference flagged for review in this section)

  • Fully stereo at the individual operator level — left and right can even run at different frequencies

  • Eight-voice unison per FM oscillator — up to 48 independently-phased operator waves per note, per FM instance

  • Operators aren't limited to a simple sine wave — any sample in the library can be used as an operator, with a smooth, band-limited morph between waveforms

Transient Oscillator

This engine is built around a classic idea from vintage digital synthesis: stapling a short, complex, sampled attack onto a sustained tone makes the ear accept the result as a single, acoustic-sounding instrument.

  • A two-operator FM pair in which both operators are full multi-sampled instruments, not simple sine waves

  • The modulator (a sampled instrument shaped by its own decay) frequency-modulates the carrier (a second sampled instrument with its own decay)

  • The ratio between the two operators spans 0.5× to 31× and can be modulated live, per channel, during the transient itself

  • Drawn from a library of 372 deep-sampled attack instruments across more than 23,000 individual audio files

  • Up to three of these can run per voice, layered over the wavetable and FM engines

🖼️ [Insert screenshot: transient oscillator controls]


Cross-Modulation

Between the different generators sits a cross-modulation bus with eight independent routes. Each route has a source (either main synthesis slot, the sub-oscillator, or any of the three sample slots), a destination, a mode, an amount, and a ratio — and every route carries its own automatic gain compensation, so the level holds steady while the tone transforms.

There are 20 modes in total, including:

  • Classic types: FM, AM, ring modulation, rectified ring mod, and phase distortion

  • Folding and index types: wavefold modulation, wavefold distortion, and phase index (which modulates read position rather than pitch)

  • Combined types: Sync+FM and Sync+FM Saw, which run hard sync and FM from the same source simultaneously

  • Spectral AM, which performs amplitude modulation in the harmonic domain, before the main oscillator renders

  • Band split, source retune, and feedback cross-mod

Two modes are unique to Phoenix:

  • Spatial modulation modulates the two stereo sides independently, from a source that is itself stereo — so the two ears can hear genuinely different modulation results from the same note, rather than the same result simply panned

  • Waveset omission works cycle-by-cycle on the waveform itself, deciding in real time which individual wave cycles play, repeat, or drop. Because it always locks to zero-crossings, the result stays musical and in tune — even though the effect can range from subtly textured to completely transformed


Modulation Sources

Phoenix has 37 modulation sources in total, feeding a matrix with no fixed slot limit.

LFOs

There's no fixed waveform list — every LFO is a hand-drawn breakpoint curve (up to 20 points), so you're never limited to a preset shape. Four rate types are available (time, frequency, sync, and a "tuned" mode where the LFO tracks note pitch and effectively becomes an oscillator). Every LFO includes Stagger and Mirror controls, applied with opposite polarity per channel — meaning an LFO routed to a stereo-aware destination produces genuine left/right divergence rather than a simple, correlated wobble.

Envelopes

Eight five-stage envelopes, each with three independently curvable segments. The curve shape itself is a continuously adjustable control (rather than a fixed set of curve types), so you can reshape a segment's character while a note is still ringing.

MetaSeq (Curve Sequencers)

Eight curve sequencers, each with up to four stages, where every stage is its own drawable curve with its own tempo-synced rate, playback mode, and trigger settings. Left and right can run entirely different sequences. Speed is bipolar, so a sequence can run backwards. By default, a MetaSeq is configured to behave like a simple envelope — but it can grow into a much more complex, evolving shape whenever you want.

The Modulation Matrix

There's no fixed number of modulation slots — routings are simply added as needed. Each routing can be:

  • Additive — layered on top of the parameter's existing value

  • Takeover — the source fully replaces the destination value, useful for handing a parameter entirely to a sequencer and taking it back later

Modulators can modulate other modulators — an LFO's speed can itself be modulated, an envelope's stage times can be pushed around by a sequencer, and so on. Every source and destination can also carry an activation delay, a tempo-syncable pre-roll of up to 4 seconds, which is itself modulatable.

Random Generators (RND)

Four random generators, each with 31 modes, split across a few families:

  • A clocked sample-and-hold, timed so it lands precisely on the beat

  • Three colored noise types (pink, red, and white) run through dedicated filters

  • Eight named stochastic (randomised) processes, each modelling a different real-world or mathematical behaviour — for example, a slow, self-correcting drift reminiscent of an analog circuit warming up, a heavy-tailed "sudden jump" process, and a Poisson-based crackle/pop generator

  • 19 rhythmic pulse-train libraries, made up of 636 hand-made rhythm patterns across styles from Euclidean rhythms to more organic, evolving grooves. Every cycle, the generator re-selects from the weighted pattern library and blends the results into a new composite rhythm — so a part can groove without ever looping identically, while still printing exactly the same "take" every time thanks to its seeded design. This is the Stutter Edit rhythmic vocabulary, rebuilt as a modulation source — and the connection runs deeper than the name: Phoenix's modulation curves convert 1:1 from Stutter Edit 2's curve format, so if you already have a library of Stutter Edit patterns, that same rhythmic vocabulary carries straight over


The Arpeggiator

A flexible step sequencer with four stages of up to 128 steps each, where every step has its own start position, length, and velocity — freely drawn rather than locked to a fixed grid. Each stage also has its own play probability, so parts of a pattern can be stochastically skipped and a sequence can slowly evolve over many bars.

On top of the raw step sequencer sit 11 melodic patterns and 8 scale modes across 12 tonics, with adjustable grid resolution and speed — so you can hold a chord and let the arpeggiator do the rest.


The Sampler & Sound Library

Every voice has three sampling slots, each able to host either a looping sample player or the transient (attack) engine described above. Sample instruments use a widely-supported instrument definition format that handles key/velocity mapping, tuning, crossfades, loop points, and round-robin sequencing.

Samples are memory-mapped rather than fully loaded into RAM, with a background process continuously keeping the right data ready — so even a very large library loads and plays without stalling the audio.

Rhythmic Looping

Above the standard sample loop sits a second, independent loop layer: a tempo-synced re-trigger with an equal-power crossfade between outgoing and incoming playback, and two independent playheads per channel. All of its parameters (delay, start offset, rate, crossfade) can be modulated. Critically, the rate can be swept all the way from a slow rhythmic re-trigger up into audio range — at which point the sample stops sounding like a loop and starts sounding like a new oscillator, with no mode switch required.


Effects

Phoenix includes 37 built-in effects, each with its own dedicated interface rather than a generic set of knobs. They're organized into a few families:

Reverbs & spaces — including a fully synthesized algorithmic reverb with no samples or impulse responses involved, a dense 8×8 feedback-matrix reverb for smooth, hi-fi tails, a comb-and-allpass reverb built to be glitched and abused, a zero-latency convolution reverb using 700 hand-made impulse responses, and two more experimental spatial/dispersion effects for shimmering, metallic, or "smeared" textures.

Delay & echo — a glitchy fracture-echo effect combined with a small reverb, a dual-engine character delay (tape-style and lo-fi digital-style lanes that can be blended), and a clean, wide stereo/ping-pong delay.

Modulation — a flexible unison/ensemble/density processor, a lush multi-voice chorus, two flangers (a smooth analog-style and an aggressive "jet plane" style), and a six-stage phaser.

Texture & granular — a granular texture engine that can turn any sound into a evolving, weather-like texture, and a detailed microsound granulator with per-grain pitch, pan, and scatter controls plus its own built-in reverb.

Character, saturation & lo-fi — a deep tape-machine emulation with true magnetic tape modelling, a warmth/saturation exciter designed for drums and bass, and a flexible multi-mode distortion with per-band control.

Rhythmic & glitch — a multiband gate that turns any sound into a rhythmic pattern, a beat-cutter that slices and rearranges audio to the grid, and a sidechain-style ducking effect with nothing needed in the sidechain input.

Low end — a pitch-tracked sub-bass generator that locks a clean sub-bass tone to whatever note is playing.

Frequency & filter — a multi-mode analog-style filter, a single-sideband frequency shifter capable of self-oscillating, metallic effects, and a five-band morphing EQ with independent left/right control.

Dynamics & mastering — a four-band upward/downward compressor, a full mastering chain (compressor, EQ, and soft clipping) in one slot, multiband and standard brickwall limiters, and a mixing utility (polarity, filtering, width, pan, and gain).

Effects are arranged in 12 global slots across three routing chains, with four routing layouts, drag-and-drop reordering, and up to four instances of the same effect running at once if needed. There are also two effect slots per individual voice, whose sends can be modulated per note and per channel — meaning each of the sixteen voices can decide, independently, how much of itself goes into which effects chain.


The Visualizer

Phoenix includes a real-time 3D display that draws the actual waveform being produced — after all the spectral and phase warping has been applied — as a shape you can rotate and zoom into. A 2D view shows the same information as a cross-section, and a dedicated FM view displays the live operator layout of the current algorithm. As modulation moves the sound, the shape on screen moves with it in real time, including any differences between the left and right channels.


Voices & Performance (MPE)

Phoenix provides 16 voices of true stereo polyphony, and is fully compatible with MPE (MIDI Polyphonic Expression) across five expression dimensions — meaning it can respond individually, per note, to strike velocity, aftertouch/pressure, vertical slide, and pitch bend when used with an MPE-capable controller. At full complexity, a single voice can be running the wavetable oscillator, an FM oscillator, three sample slots, a sub-oscillator, eight cross-modulation routes, two filters, and all 37 modulation sources simultaneously — multiplied by up to sixteen notes at once.


Getting More Help

If you run into an issue that isn't covered here, reach out to support and we'll be happy to help.

Phoenix is a collaboration between BT and Spitfire Audio.


Did this answer your question?