正在自动登录…
← 返回项目列表mw00/
mw00/ project-maya
Project Maya 是一个在自有 NVIDIA GPU 上运行 GLM-5.3-Flash 大模型的引擎、服务器和仪表盘,支持聊天、图片识别和兼容 OpenAI/Anthropic 的 API,适合希望在本地部署大型模型的个人或团队。
AI 中文解读 · 根据仓库描述与 README 生成 · 2026/10/11 14:06:17 · 查看原文
- 累计 Stars
- 321
- 当前热度
- 1~7 天 · 小火
- 仓库创建时间
选用指南
AI 根据 README 整理 · 2026/10/11 14:06:17
- 适用场景
适合在本地或自有 GPU 服务器上部署大型语言模型,用于聊天、图片识别和集成到其他应用或编码代理。
查看 README 依据
<h1 align="center">Project Maya</h1> <p align="center"><b>Run GLM-5.3-Flash - a 321-billion-parameter AI model - on your own GPU(s)</b><br> One NVIDIA GPU or several (up to 16), AMD (experimental) · Linux, Windows (experimental) · chat in the browser, pictures, OpenAI- and Anthropic-compatible API</p> <p align="center"><a href="https://buymeacoffee.com/peasantsmith">☕ Support Project Maya - buy me a coffee</a></p> Maya runs **[GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash)** (zai-org, MIT license): a mixture-of-
edit and regenerate, code with a preview for HTML pages), a live **Monitor** of the model, the expert caches and your GPU/CPU/RAM (with **Copy report** for an issue), and **About** (the version and its **Update** button when a new one is out, this PC, how to connect your tools). **Settings** (the gear in the chat) changes the **context size**: the model reloads with it in a minute or two, and the dashboard shows what the size costs in GPU memory on this PC. - **Your apps and coding agents:** an "OpenAI-compatible" provider with the base URL `http://127.0.0.1:8080/v1`
- 使用方式
通过命令行安装和运行,提供浏览器仪表盘和 OpenAI/Anthropic 兼容 API。
查看 README 依据
(experimental, Linux with ROCm 7, or Windows): `./maya.sh --backend hip --gpu 0 --check` first (Windows: `START-MAYA.bat --backend hip --gpu 0 --check`; two cards: `--gpus 0,1`; Strix Halo stays one GPU), then [docs/AMD_MAYA.md](docs/AMD_MAYA.md). 1. Get Project Maya: ```sh git clone https://github.com/mw00/project-maya.git && cd project-maya ``` (or [download it](https://github.com/mw00/project-maya/archive/refs/heads/main.zip) and unzip it). 2. Run **`./setup.sh`** (the same as `./maya.sh`).
| `--calibrate` | tune the engine's CPU lane for this PC (see [Tuning](#tuning)), then start; with `--no-start` only tune | | `--yes` | the recommended answers (the model download still needs `--download-model`) | | `--plain` | the setup in plain text with typed answers, and Maya in the terminal, not on their own screen (a pipe gets this too, and `--yes` a plain setup); the setup screen's whole output is in `maya-setup.log` | ## Using it - **In the browser:** `http://127.0.0.1:8080` - **Chat** (your chats kept in this browser, instructions per chat,
edit and regenerate, code with a preview for HTML pages), a live **Monitor** of the model, the expert caches and your GPU/CPU/RAM (with **Copy report** for an issue), and **About** (the version and its **Update** button when a new one is out, this PC, how to connect your tools). **Settings** (the gear in the chat) changes the **context size**: the model reloads with it in a minute or two, and the dashboard shows what the size costs in GPU memory on this PC. - **Your apps and coding agents:** an "OpenAI-compatible" provider with the base URL `http://127.0.0.1:8080/v1`
- 部署要求
需要 NVIDIA GPU(计算能力 7.0 或更高,如 V100 或更新),Linux 系统(Windows 实验性支持),约 100GB 可用 NVMe SSD 空间,以及 NVIDIA 驱动和 CUDA 工具包 12.x。
查看 README 依据
| **GPU** | NVIDIA, compute capability 7.0 or newer (V100 and newer); one GPU, or up to 16 that share the model (two split the layers in the middle; with more, each takes a share sized to its free VRAM). The engine fills whatever VRAM you have with the most-used experts: more VRAM is faster. Measured: 1 and 2x V100 32 GB; by users: 2x CMP 170HX (above), 1x RTX 3090 and nine GPUs (8x RTX 5060 Ti 16 GB + the 3090). **AMD (experimental):** RX 7900 XT / XTX, Radeon AI PRO R9700 / RX 9070 (one GPU or two) and Strix Halo / Gorgon Halo, Radeon 8060S / 8065S (one GPU), text only
| **System** | Linux (x86-64; a CPU with AVX2 is best - without it the engine still runs, its CPU expert lane on ggml's slower kernels), NVIDIA driver, CUDA toolkit 12.x (CUDA 13 can be used for Turing and newer, but it no longer compiles for Volta/V100), g++, Python 3.10+. Windows 10/11: experimental, with Visual Studio 2022 Build Tools instead of g++ ([Windows](#windows)). Not WSL2. AMD: ROCm 7 instead of the NVIDIA driver and CUDA (on Windows AMD's ROCm SDK wheels, which the setup offers to install). |
The installer checks all of this and prints the exact command for anything missing. It installs nothing system-wide by itself. ## Install **You need:** an NVIDIA GPU (V100 / RTX 20 or newer) on Linux (Windows: [experimental](#windows)), ~100 GB free on an NVMe SSD, a current NVIDIA driver and the CUDA toolkit (12.x for a V100; the engine is compiled for your GPU). Everything else - Python, the engine, the model - is set up for you, the way Strata does it. On an AMD RX 7900 XT / XTX, R9700 / RX 9070 or Strix Halo / Gorgon Halo
- 已知限制
AMD GPU 支持为实验性,且仅支持文本;Windows 支持为实验性;不支持 WSL2。
查看 README 依据
for GLM-5.3-Flash (the expert tiers across VRAM, RAM and SSD, the two-GPU split, MTP decoding), and its server and dashboard started from Strata's and were reworked for Maya (a new dashboard, images on demand, the thinking budget). **AMD (experimental):** Linux and Windows on RX 7900 XT / XTX, R9700 / RX 9070 and Strix Halo / Gorgon Halo (Radeon 8060S / 8065S), one GPU or two (Strix Halo: one), text only - see [docs/AMD_MAYA.md](docs/AMD_MAYA.md). ## The models: Maya-S, Maya-S24, Maya-M, Maya-M-Derisked and Maya-L
| **System** | Linux (x86-64; a CPU with AVX2 is best - without it the engine still runs, its CPU expert lane on ggml's slower kernels), NVIDIA driver, CUDA toolkit 12.x (CUDA 13 can be used for Turing and newer, but it no longer compiles for Volta/V100), g++, Python 3.10+. Windows 10/11: experimental, with Visual Studio 2022 Build Tools instead of g++ ([Windows](#windows)). Not WSL2. AMD: ROCm 7 instead of the NVIDIA driver and CUDA (on Windows AMD's ROCm SDK wheels, which the setup offers to install). |
README
正在加载 README…
收录信息
- 仓库创建时间
- 2026/10/7 06:55:13
- 首次达标收录
- 2026/10/11 11:47:02
- 统计更新时间
- 2026/10/11 15:35:43
按创建时长与累计 Stars 判断等级。首次达标收录后,即使未达当前门槛仍保留;累计 Stars 和仓库状态会定期更新。
模型评估
值得看加权总分 4.0 / 5
由 Jev 模型根据仓库描述与 README 评估,2026/10/11 11:59:36 生成。评分仅供参考,不是代码、安全或许可证审查;热度等级另按创建时长与累计 Stars 计算。
仓库信息
- 最近推送
- 2026/10/11 13:10:37
- 许可证标识
- MIT
- Fork 数
- 45
- 当前状态
- 未归档
收录记录
按创建时长与累计 Stars 首次收录:小火
首次发现项目