• 0 Posts
  • 89 Comments
Joined 2 years ago
cake
Cake day: June 2nd, 2024

help-circle




  • If you’re coding up a custom system, Try adding a 9b model that dose RAG searches with chromadb and outputs a memory summary that feeds into the main models context window. You can ingest old messages, web pages, character info, or whatever text. Qwen 3.6 35b is the first local model I ran that seems to be good enough to be useful if it has some infrastructure around it, and smart enough to not die on a tool call(usually).

    Add in DDGS search and a web page scrape tool that adds to the RAG system, let it loop a few times and its actually quite useful, even if it takes 5-10 minutes to get an output. After that it’s all tweaking what is and isn’t in in context.

    Also, try not to go too insane.





  • Running decencored Qwen3.6-27b and a 9b Gemma for RAG and scrapes on Ollama with a mostly vibe coded discord bot. Just got it to run tools and scrape and post news on a schedule. The first model I can run locally that’s smart enough to be useful. May give Jan a try for the back end after reading that other guys rant.

    Mostly use it for stupid questions I could have googled and to brag to friends.