Aurora Both Mode: Separate System Audio and Microphone on Windows

Most captioning sessions have one obvious source: a meeting microphone or the audio coming from a Windows app. Some workflows need both at once. Aurora Subtitles includes Both (experimental) for that case. It separates system audio and microphone capture so the two inputs can be handled as distinct parts of the session rather than treated as one undifferentiated stream.
What the mode is for
Both mode is useful when you need to follow what another person says while also monitoring your own microphone activity. Examples include a multilingual call, a remote class, a game session with voice chat, or a stream where the PC audio and the presenter’s microphone have different roles.
The product documentation makes an important distinction: microphone text stays in Session Activity instead of appearing in the main overlay. System audio can continue to provide the visible caption layer, while microphone events remain available in the activity view. This is a workflow choice, not a claim that every captured word will always be displayed in the same place.
Aurora also documents experimental anonymous speaker distinction using CPU only and no additional VRAM. Treat that as an aid for inspecting a session, not as a promise of perfect speaker identification. Accuracy can change with overlapping speech, background noise, accents, microphone placement, and the selected language route.
A controlled setup
- Confirm the basics in the requirements and policy section: Windows 11 64-bit, enough disk space for the selected local models, and internet access for activation or model downloads when required.
- Review the live demo so you know which area is the overlay and which area represents session activity.
- Choose the system-audio and microphone inputs deliberately. Aurora uses WASAPI for Windows system audio and microphone capture; avoid changing the Windows default device halfway through the first test.
- Start with a short recording or low-stakes call. Check that system audio produces the visible captions you expect and that microphone events are kept in Session Activity rather than the overlay.
- Only then add translation or experimental speaker distinction. Isolating one change at a time makes it easier to tell whether a problem comes from input routing, language selection, audio quality, or overlay readability.
The overlay itself remains configurable: placement, width, height, font, color, outline, background, fade time, and line count can be adjusted for the app underneath. Click-through behavior is useful when the overlay sits above a game or meeting window and should not intercept normal clicks.
Trade-offs to understand
Both mode is marked experimental, so it is best for users who value separated inputs and are willing to validate the result in their own hardware and audio setup. It does not turn Aurora into a medical device, a certified interpreter, or a replacement for human review. For accessibility workflows, it can add a useful caption layer, but important conversations and decisions still deserve an appropriate professional or human check.
The commercial model remains simple: Aurora is a Windows desktop product with a one-time payment and no monthly subscription or per-minute cloud credit model. Review the product and checkout section before purchasing. If you already own a valid license, use the download page for the current installer rather than treating this guide as a free trial offer.
If Both mode is more than you need, start with Maximum Compatibility for CPU-only live captions. It is the simpler path for English input audio on a PC without a dedicated GPU; the linked guide explains its 51 published translation targets.