Structured framework for evaluating LLM outputs using Google Gemini, applying software testing principles to AI validation