Finance FAQ Assistant

Ask a personal-finance question and compare the answer from the base model (untrained Qwen2.5-0.5B) against the fine-tuned model (non-instruction domain adaptation → instruction SFT → DPO alignment) side by side.

Base model (before fine-tuning)

Fine-tuned model (after DPO alignment)

Examples