A Jev-like wrapper for LLMs, including vision models
Summary
Allan Riordan Boll presents a Jev-like wrapper for LLMs with multimodal support, including vision models, and demonstrates reading token logprobs to select answers. The post includes a standalone Python example, image attachments, and a practical comparison between Gemma-4-12B on an RTX 3090 and OpenAI services, highlighting cost and latency considerations. It emphasizes flexibility in prompt-driven reasoning and local experimentation.