Skip to content

Food for thought #1

Description

@Wladastic

I am thinking about renaming the project to make it a bit more descriptive and trustworthy.

I will keep using "AI" helpers like GitHub Copilot to continue developing it and want to have at least one project in my repository that's public and shows that I'm actually building something, not just freelancing or working somewhere.

Since this project is also meant to later work with Chat.fans agent task calls, I might turn it into a more general MCP browser automation tool.

I've recently added support for screenshot analysis using Gemma3:4b together with a pretext parameter. That way, the LLM calling the screenshot function can give a bit more context on what it's supposed to look for. But I'm still not really happy with the results. Gemma3:4b often just asks what it should do next with the picture instead of actually giving a clean answer or even gives suggestions about what it sees and offers help improving it, which confused even Claude Sonnet 4 into hallucinating that it was talking back and forth with it.

Maybe changing the prompt would help but so far I haven't seen models under 12B from Gemma follow instructions very well.
So let's wait for Qwen3 VL models to show up.
Qwen3:4b for example does follow instructions quite well although its German is sometimes a bit broken or it mixes things up in context. Like when a story says Person A said X to Person B and it suddenly turns it around and says Person A is being told X by Person B which breaks the narrative.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions