Comments (3)
@eduardozamudio can you please share the version of vLLM used and the error of why the model won't load? It is a MistralForCausalLM
, so I would expect it to run as Mistral models do.
from vllm.
Hi @eduardozamudio (Hola Eduardo!)
mistralai/Codestral-22B-v0.1
worked for me using vllm 0.4.3.
Regards
Matias
from vllm.
Hi @eduardozamudio (Hola Eduardo!)
mistralai/Codestral-22B-v0.1
worked for me using vllm 0.4.3.Regards Matias
Does it work for "fill in the middle" https://huggingface.co/mistralai/Codestral-22B-v0.1#fill-in-the-middle-fim? I image there would be some work required to support the prefix and suffix params both in the rest API and core APIs...
I haven't dug into the https://github.com/mistralai/mistral-inference code yet, but I think it's just uses special tokens to mark prefix, suffix and middle, so it can probably also be implemented outside of vllm and just passed as the normal input...
from vllm.
Related Issues (20)
- [Bug]: RuntimeError: CUDA error: no kernel image is available for execution on the device CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
- [Speculative decoding]: The content generated by speculative decoding is inconsistent with the content generated by the target model HOT 3
- [Performance]: Is kv cache implemented globally in vllm, It can be shared by multiple concurrent inferences? HOT 2
- [Feature]: Make a unstable latest docker image HOT 1
- [Performance]: gptq and awq quantization do not improve the performance HOT 4
- [Installation]: Compiling VLLM for cpu only. HOT 2
- [Usage]: Streaming Response from vLLM 0.4.2 -> 0.4.3 HOT 2
- [Usage]: Function calling for mistral v0.3
- [Bug]: vLLM does not support virtual GPU
- [Bug]: Unexpected prompt token logprob behaviors of llama 2 when setting echo=True for openai-api server HOT 1
- [Bug]: Getting an empty string ('') for every call on fine-tuned Code-Llama-7b-hf model HOT 1
- [Bug]: non-deterministic Python gc order leads to flaky tests HOT 13
- [Usage]: Howto quiet the terminal 'Info' outputs in vllm HOT 1
- [Performance]: [Automatic Prefix Caching] When hitting the KV cached blocks, the first execute is slow, and then is fast. HOT 3
- [Speculative decoding]: `AttributeError: 'NoneType' object has no attribute 'numel'` when exceeding draft context length HOT 3
- [Bug]: Qwen2 MoE: AttributeError: 'MergedColumnParallelLinear' object has no attribute 'weight'. Did you mean: 'qweight'? HOT 3
- [Bug]: with `--enable-prefix-caching` , `/completions` crashes server with `echo=True` above certain prompt length HOT 2
- [RFC]: Refactor MoE
- [Bug]: TorchSDPAMetadata is out of date HOT 2
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from vllm.