lambda-1-160m-base

lambda-1-160m-base is an experimental language model created with a custom myllm decoder-only Transformer implementation.

All training code is publicly available at KeisukeMiyamoto1324/myllm.

Model Details

Item Value
Parameters 164.5M
Architecture Decoder-only Transformer
Context length 1024 tokens
Tokenizer Byte-level BPE
Vocabulary size 65,536
Layers 16
Hidden size 768
Attention heads 12
FFN size 3,072

Training Data

The model was pretrained on a Japanese text mixture.

Dataset Notes
MK0727/CleanedFineWeb2Edu-jp Filtered Japanese web corpus
MK0727/SyntheticTextbook-jp Synthetic Japanese corpus

Usage

git clone https://github.com/KeisukeMiyamoto1324/lambda.git
cd lambda
python3 -m venv venv
source venv/bin/activate
pip3 install -r requirements.txt

python3 src/inference_base/inference_hf.py \
  --prompt "人工知能とは" \
  --max-new-tokens 64

Limitations

This model is not instruction-tuned or safety-aligned. It may generate incorrect, biased, unsafe, or low-quality text.

The model was trained on a limited Japanese corpus mixture and has not been evaluated on standard benchmarks.


Support Lambda

Lambda is an open-source project for building small Japanese language models from scratch. As a student, I have funded this project with income from my part-time job, but the growing training costs are becoming difficult to cover.

Your support helps cover GPU costs and develop larger models. Thank you for helping Lambda continue to grow.

Vast.ai

Vast.ai offers affordable cloud GPUs for AI training, with NVIDIA H100 SXM GPUs available from around $1.54 per hour. If you purchase credits through the link below, I receive 3% in GPU credits at no extra cost to you.

https://cloud.vast.ai/?ref_id=521936

Ko-fi

Support Lambda with a donation starting from $5.

Support Lambda on Ko-fi
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support