Instructions to use replit/replit-code-v1-3b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use replit/replit-code-v1-3b with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="replit/replit-code-v1-3b", trust_remote_code=True)

# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("replit/replit-code-v1-3b", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("replit/replit-code-v1-3b", trust_remote_code=True)

Notebooks
Google Colab
Kaggle
Local Apps

vLLM

How to use replit/replit-code-v1-3b with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "replit/replit-code-v1-3b"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "replit/replit-code-v1-3b",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Use Docker

docker model run hf.co/replit/replit-code-v1-3b

SGLang

How to use replit/replit-code-v1-3b with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "replit/replit-code-v1-3b" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "replit/replit-code-v1-3b",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "replit/replit-code-v1-3b" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "replit/replit-code-v1-3b",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Docker Model Runner
How to use replit/replit-code-v1-3b with Docker Model Runner:
```
docker model run hf.co/replit/replit-code-v1-3b
```

Speed UP method

#14

by luoji12345 - opened May 19, 2023

Discussion

luoji12345

May 19, 2023

Replit is an amazing model as it can generate valid results even thouth it is only 2.7B.
However I have trouble in accelerating Inference Replit. These are the method that I have ever tried and failed :

torch2.0 : is an easy way to realize, but no speed up than torch 1.13
flash_attn: not match alibi , so no model weights available
triton : I meed the same error as described in https://huggingface.co/mosaicml/mpt-7b-storywriter/discussions/10
deepspeed : Realized but no speed up than torch 1.13, and it only fits torch.float16 rather than torch.bfloat16, slower in A100

I am going to try FasterTransformer although it is hard to realize.

Could you please give me some advice about the inference accelerate ? Will you release the project to accelerate the Repilt Inference ?

leojames

May 24, 2023

I used the NVIDIA GPU P100, which is based on the Pascal architecture, so it does not support this type of acceleration. Even after switching to A30, the speed is still slow.I donot know how to accerlerate the repilt interfence .Offical can give some advice ？

pirroh

Replit org Jun 5, 2023

Closing issue as OP has already found a solution and posted it on the mpt-7b-storywriter discussion thread: https://huggingface.co/mosaicml/mpt-7b-storywriter/discussions/10#646833123a7c8dda230f87ab

pirroh changed discussion status to closed Jun 5, 2023

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment