We can run LLM models locally, assuming we have enough VRAM or a big enough MAC, and use those models for Coding assist. No messages have to leave our network and we don't have to worry about unexpected costs due to orphan processes. This can also work for local servers
Windows with NVidia 8GB: https://joe.blog.freemansoft.com/2024...
Mac: https://joe.blog.freemansoft.com/2024...
Windows with NVidia 24GB: https://joe.blog.freemansoft.com/2024...