Finance FAQ Assistant
Ask a personal-finance question and compare the answer from the base model (untrained Qwen2.5-0.5B) against the fine-tuned model (non-instruction domain adaptation → instruction SFT → DPO alignment) side by side.
Base model (before fine-tuning)
Fine-tuned model (after DPO alignment)
Examples