GBP/USD
EUR/USD
USD/JPY
GBP/EUR
Gold $/oz
Silver $/oz

Tech

A User Managed to Run 27B Language Models by Modifying the Nvidia Tesla V100 Graphics Card Into a Gaming Computer

News summary

  • Details are in our news.
  • Running high-quality language models requires the graphics card to have large video memory (VRAM).
  • An RTX 4080 owner realized that his system, which can easily play modern games, was insufficient for large language models (LLM) and made an interesting modification.
  • The user managed to achieve both gaming and artificial intelligence performance by integrating the NVIDIA Tesla V100 graphics card into his gaming computer.

A user managed to run 27B language models by modifying the NVIDIA Tesla V100 graphics card into a gaming computer. Details are in our news. Running high-quality language models requires the graphics card to have large video memory (VRAM).

An RTX 4080 owner realized that his system, which can easily play modern games, was insufficient for large language models (LLM) and made an interesting modification. The user managed to achieve both gaming and artificial intelligence performance by integrating the NVIDIA Tesla V100 graphics card into his gaming computer. The technical difficulties encountered in this process were overcome with special adapters and cooling solutions.

Connecting the Tesla V100 board to a standard desktop motherboard required the use of an SXM2-PCIe converter adapter. This installation, which cost approximately $ 266 in total, enabled the card with 16GB HBM2 memory to be included in the system. Tesla V100 stands out with its 5,120 CUDA cores and 4,096-bit bus offering 900GB/s bandwidth.

The lack of a video output or standard PCIe power connection on the card required additional effort for the modder during the installation phase. The high noise level of 82dB produced by the cooling system was controlled by using a 9V battery and PWM jumper. The fan speed was reduced to 10 percent of its original maximum value, ensuring quieter operation of the system.

The 32GB total VRAM added to the system provided the necessary space to run models such as the Qwen3.6 27B. The model, which occupies 19GB at the Q5_K_M quantization level, can be run at a speed of 32 tokens per second with a context size of 128K tokens. The transaction processing rate varies between 133 and 160 tokens per second. This setup offers a local AI work environment that doesn't require an internet connection for under $300.

Source: ShiftDelete

Most read in this category