> ## Documentation Index
> Fetch the complete documentation index at: https://dragonwingdocs.qualcomm.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Whisper

Whisper 是 OpenAI 的通用自动语音识别(ASR)模型。您可以将其用于音频转录、翻译和语言识别。您可以使用 Qualcomm 的 [VoiceAI ASR](https://softwarecenter.qualcomm.com/catalog/item/VoiceAI_ASR_Community) 在 Dragonwing 开发板的 NPU 上运行 Whisper,或使用 [whisper.cpp](https://github.com/ggml-org/whisper.cpp) 在 CPU 上运行。

## 使用 VoiceAI ASR 在 NPU 上运行 Whisper

### 1. 安装 SDK

1. 在开发板上打开终端,并为此示例设置基本依赖:

   ```
   sudo apt install -y cmake pulseaudio-utils
   ```

2. 安装 [AI Runtime SDK - Community Edition](https://softwarecenter.qualcomm.com/catalog/item/Qualcomm_AI_Runtime_Community):

   ```
   # Install the SDK
   wget -qO- https://cdn.edgeimpulse.com/qc-ai-docs/device-setup/install_ai_runtime_sdk.sh | bash

   # Use the SDK in your current session
   source ~/.bash_profile
   ```

3. 安装 [VoiceAI ASR - Community Edition](https://softwarecenter.qualcomm.com/catalog/item/VoiceAI_ASR_Community):

   ```shell theme={null}
   cd ~/

   wget https://softwarecenter.qualcomm.com/api/download/software/sdks/VoiceAI_ASR_Community/All/2.3.0.0/VoiceAI_ASR_Community_v2.3.0.0.zip
   unzip VoiceAI_ASR_Community_v2.3.0.0.zip -d voiceai_asr
   rm VoiceAI_ASR_Community_v2.3.0.0.zip

   cd voiceai_asr/2.3.0.0/

   # Put the path to VoiceAI ASR in your bash_profile (so it's available under VOICEAI_ROOT)
   echo "" >> ~/.bash_profile
   echo "# Begin VoiceAI ASR" >> ~/.bash_profile
   echo "export VOICEAI_ROOT=$PWD" >> ~/.bash_profile
   echo "# End VoiceAI ASR" >> ~/.bash_profile
   echo "" >> ~/.bash_profile

   # Re-load the environment variables
   source ~/.bash_profile

   # Symlink Whisper libraries
   cd $VOICEAI_ROOT/whisper_sdk/libs/npu/rpc_libraries/linux/whisper_all_quantized/
   sudo ln -s $PWD/*.so /usr/lib/
   ```

### 2. 从 AI Hub 下载模型

安装好 SDK 后,您可以从 [AI Hub](https://aihub.qualcomm.com/models?searchTerm=whisper) 下载预编译的 Whisper 模型。下载模型时,请选择以下设备:

* RB3 Gen 2 Vision Kit:'Qualcomm QCS6490 (Proxy)'
* RUBIK Pi 3:'Qualcomm QCS6490 (Proxy)'
* IQ-9075 EVK:'Qualcomm QCS9075 (Proxy)'

下载后,将编码器模型重命名为 `encoder_model_htp.bin`,将解码器模型重命名为 `decoder_model_htp.bin`。

要在开发板上直接下载 [Whisper-Small-Quantized](https://aihub.qualcomm.com/models/whisper_small_quantized?searchTerm=whisper\&chipsets=qualcomm-qcs6490-proxy) 模型:

* RB3 Gen 2 Vision Kit / Rubik Pi 3:

  ```shell theme={null}
  mkdir -p ~/whisper_models/ai_hub_small_quantized/
  cd ~/whisper_models/ai_hub_small_quantized/

  # Models from https://aihub.qualcomm.com/models/whisper_small_quantized for QCS6490 (Proxy) target
  wget -O encoder_model_htp.bin https://huggingface.co/qualcomm/Whisper-Small-Quantized/resolve/0e21411/precompiled/qualcomm-qcs6490-proxy/Whisper-Small-Quantized_WhisperSmallEncoderQuantizable_w8a16.bin
  wget -O decoder_model_htp.bin https://huggingface.co/qualcomm/Whisper-Small-Quantized/resolve/0e21411/precompiled/qualcomm-qcs6490-proxy/Whisper-Small-Quantized_WhisperSmallDecoderQuantizable_w8a16.bin

  # Vocab file is not in AI Hub yet, grab from our CDN
  wget -O vocab.bin https://cdn.edgeimpulse.com/qc-ai-docs/models/whisper/vocab.bin
  ```

* IQ-9075 EVK:

  ```shell theme={null}
  mkdir -p ~/whisper_models/ai_hub_small_quantized/
  cd ~/whisper_models/ai_hub_small_quantized/

  # Models from https://aihub.qualcomm.com/models/whisper_small_quantized for QCS9075 (Proxy) target
  wget -O encoder_model_htp.bin https://huggingface.co/qualcomm/Whisper-Small-Quantized/resolve/0e21411/precompiled/qualcomm-qcs9075-proxy/Whisper-Small-Quantized_WhisperSmallEncoderQuantizable_w8a16.bin
  wget -O decoder_model_htp.bin https://huggingface.co/qualcomm/Whisper-Small-Quantized/resolve/0e21411/precompiled/qualcomm-qcs9075-proxy/Whisper-Small-Quantized_WhisperSmallDecoderQuantizable_w8a16.bin

  # Vocab file is not in AI Hub yet, grab from our CDN
  wget -O vocab.bin https://cdn.edgeimpulse.com/qc-ai-docs/models/whisper/vocab.bin
  ```

### 3. 编译并运行示例

1. 构建 `npu_rpc_linux_sample/voice-ai-ref` 示例:

   ```
   cd $VOICEAI_ROOT/whisper_sdk/sampleapp/npu_rpc_linux_sample/voice-ai-ref

   # Overwrite the LogUtil.h function to log to stdout
   wget -O include/LogUtil.h https://cdn.edgeimpulse.com/qc-ai-docs/code/voiceai_asr/2.3.0.0/whisper_sdk/sampleapp/npu_rpc_linux_sample/voice-ai-ref/include/LogUtil.h
   # Overwrite the main.cpp example to add microphone selection
   wget -O src/main.cpp https://cdn.edgeimpulse.com/qc-ai-docs/code/voiceai_asr/2.3.0.0/whisper_sdk/sampleapp/npu_rpc_linux_sample/voice-ai-ref/src/main.cpp

   # Symlink Whisper libraries for build
   mkdir -p libs/arm64-v8a/
   cd libs/arm64-v8a/
   ln -s $VOICEAI_ROOT/whisper_sdk/libs/npu/rpc_libraries/linux/whisper_all_quantized/*.so .
   cd ../../

   mkdir -p build
   cd build
   cmake ..
   make -j`nproc`
   ```

2. 现在您可以转录 .WAV 文件:

   ```
   cd $VOICEAI_ROOT/whisper_sdk/sampleapp/npu_rpc_linux_sample/voice-ai-ref/build

   # Download sample file
   wget -O jfk.wav https://raw.githubusercontent.com/ggml-org/whisper.cpp/refs/heads/master/samples/jfk.wav

   # Transcribe:
   ./voice-ai-ref -f jfk.wav -l en -t transcribe -m ~/whisper_models/ai_hub_small_quantized/

   # ... Expected result:
   # VoiceAIRef final result =  And so my fellow Americans, ask not what your country can do for you, ask what you can do for your country. [language: English]

   # Press Control-C to exit the running application.
   ```

3. 甚至可以进行实时转录:

   a. 将麦克风连接到开发板。

   b. 找到麦克风的名称:

   ```
   pactl list short sources
   # 49	alsa_output.platform-sound.stereo-fallback.monitor	PipeWire	s24-32le 2ch 48000Hz	SUSPENDED
   # 76	alsa_input.usb-046d_C922_Pro_Stream_Webcam_C72F6EDF-02.analog-stereo	PipeWire	s16le 2ch 32000Hz	SUSPENDED

   # To use the USB webcam, use "alsa_input.usb-046d_C922_Pro_Stream_Webcam_C72F6EDF-02.analog-stereo" as the name
   ```

   3. 运行实时转录:

      ```
      ./voice-ai-ref -r -l en -t transcribe -m ~/whisper_models/ai_hub_small_quantized/ -d "alsa_input.usb-046d_C922_Pro_Stream_Webcam_C72F6EDF-02.analog-stereo"

      # VoiceAIRef final result =  Hi, this is to see if I can do live transcription. [language: English]
      ```

<Danger>当 VAD 判断没有语音后,实时转录会立即报错退出,希望这一问题会在未来的更新中得到修复。</Danger>

🚀 您现在可以在开发板上完全离线地转录音频了!VoiceAI ASR 没有面向高级语言(如 Python)的绑定,因此如果您想在应用程序中使用 Whisper,最简单的方法是直接启动 `voice-ai-ref` 二进制文件,并从 `stdout` 读取数据。

## 使用 whisper.cpp 在 CPU 上运行 Whisper

您也可以使用 whisper.cpp(或其他任何流行的 Whisper 库)在 CPU 上运行 Whisper(性能较低)。

以下是 [whisper.cpp](https://github.com/ggml-org/whisper.cpp) 的操作说明。在开发板上打开终端,或通过 SSH 会话连接到开发板,然后运行:

1. 安装构建依赖项:

   ```
   sudo apt update
   sudo apt install -y libsdl2-dev libsdl2-2.0-0 libasound2-dev
   ```

2. 构建 whisper.cpp:

   ```
   mkdir -p ~/dev/llm/
   cd ~/dev/llm/

   git clone https://github.com/ggml-org/whisper.cpp.git
   cd whisper.cpp
   git checkout v1.7.6

   # Build (CPU)
   cmake -B build-cpu -DWHISPER_SDL2=ON
   cmake --build build-cpu -j`nproc` --config Release
   ```

3. 将 whisper.cpp 路径添加到您的 PATH:

   ```
   cd ~/dev/llm/whisper.cpp/build-cpu/bin

   echo "" >> ~/.bash_profile
   echo "# Begin whisper.cpp" >> ~/.bash_profile
   echo "export PATH=\$PATH:$PWD" >> ~/.bash_profile
   echo "# End whisper.cpp" >> ~/.bash_profile
   echo "" >> ~/.bash_profile

   # To use the whisper.cpp files in your current session
   source ~/.bash_profile
   ```

4. 现在您可以使用 whisper.cpp 转录一些音频:

   ```
   # Download model
   cd ~/dev/llm/whisper.cpp
   sh ./models/download-ggml-model.sh tiny.en-q5_1

   # Transcribe text
   whisper-cli -m models/ggml-tiny.en-q5_1.bin -f samples/jfk.wav

   # [00:00:00.000 --> 00:00:10.480]
   # and so my fellow Americans ask not what your country can do for you ask what you can do for your country
   ```

5. 您还可以实时转录音频:

   a. 将麦克风连接到开发板。

   b. 找到麦克风 ID:

   ```
   SDL_AUDIODRIVER=alsa whisper-stream -m models/ggml-tiny.en-q5_1.bin
   # init: found 2 capture devices:
   # init:    - Capture device #0: 'qcm6490-rb3-vision-snd-card, '
   # init:    - Capture device #1: 'Yeti Stereo Microphone, USB Audio'

   # If you want "Yeti Stereo Microphone, USB Audio" then the ID is 1
   ```

   c. 开始实时转录:

   ```
   SDL_AUDIODRIVER=alsa whisper-stream -m models/ggml-tiny.en-q5_1.bin -c 1

   # main: processing 48000 samples (step = 3.0 sec / len = 10.0 sec / keep = 0.2 sec), 4 threads, lang = en, task = transcribe, timestamps = 0 ...
   # main: n_new_line = 2, no_context = 1
   #
   # [Start speaking]
   # This is a test to see if you can transcribe text live on your Qualcomm device
   ```

### 使用 OpenCL 在 GPU 上运行

您还可以构建在 GPU 上运行的二进制文件:

1. 首先按照 [llama.cpp](/zh/ai-workflows/llama-cpp) 中"安装 OpenCL 头文件和 ICD 加载器库"下的步骤操作。

2. 使用 OpenCL 构建二进制文件:

   ```
   cd ~/dev/llm/whisper.cpp

   cmake -B build-gpu -DGGML_OPENCL=ON  -DWHISPER_SDL2=ON
   cmake --build build-gpu -j`nproc` --config Release

   # Find the binary in:
   #     build-gpu/bin/whisper-cli
   ```
