cygon ,

Tip: try Oobabooga's Text Generation WebUI with one of the WizardLM Uncensored models from HuggingFace in GGML or GGUF format.

The GGML and GGUF formats perform very well with CPU inference when using LLamaCPP as the engine. My 10 years old 2.8 GHz CPUs generate about 2 words per second. Slightly below reading speed, but pretty solid. Just make sure to keep to the 7B models if you have 16 GiB of memory and 13B models if you have 32 GiB of memory.

  • All
  • Subscribed
  • Moderated
  • Favorites
  • random
  • chatgpt@lemmy.world
  • test
  • worldmews
  • mews
  • All magazines