Creative datasets connecting Japan, Iceland, mathematics, technology, philosophy, Asian culture and Christian ethics.
Guilherme Monteiro
guicybercode
·
AI & ML interests
smol-llms, llm-training, chinese-history, brazilian-history, production-ml, terminal-ui, data-curation, edge-ai
Recent Activity
updated a collection 1 day ago
Culture, Mathematics, Technology & Ethics updated a collection 1 day ago
Culture, Mathematics, Technology & Ethics updated a collection 1 day ago
Culture, Mathematics, Technology & EthicsOrganizations
None yet
LLM Training from Scratch
Resources, datasets and base models for training small LLMs on consumer hardware. Focus on pre-training pipelines and data curation.
Smol LLMs for History
Small language models (≤3B params) that can be fine-tuned or run locally for history-focused tasks. Curated for training on Chinese and Brazilian hist
-
HuggingFaceTB/SmolLM2-135M
Text Generation • 0.1B • Updated • 2.46M • 228 -
HuggingFaceTB/SmolLM2-135M-Instruct
Text Generation • 0.1B • Updated • 1.41M • 405 -
HuggingFaceTB/SmolLM2-360M
Text Generation • 0.4B • Updated • 361k • 124 -
HuggingFaceTB/SmolLM2-360M-Instruct
Text Generation • 0.4B • Updated • 286k • 213
Papers I'm Reading
Key papers on small LLMs, data-centric training, and multilingual NLP.
-
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Paper • 2406.17557 • Published • 106 -
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Paper • 2502.02737 • Published • 261 -
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Paper • 2506.05209 • Published • 65
Multilingual models
Multilingual models featuring Portuguese, Chinese, and other languages essential for cross-cultural historical research. Curated for fine-tuning on Ch
Culture, Mathematics, Technology & Ethics
Creative datasets connecting Japan, Iceland, mathematics, technology, philosophy, Asian culture and Christian ethics.
Papers I'm Reading
Key papers on small LLMs, data-centric training, and multilingual NLP.
-
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Paper • 2406.17557 • Published • 106 -
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Paper • 2502.02737 • Published • 261 -
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Paper • 2506.05209 • Published • 65
LLM Training from Scratch
Resources, datasets and base models for training small LLMs on consumer hardware. Focus on pre-training pipelines and data curation.
Multilingual models
Multilingual models featuring Portuguese, Chinese, and other languages essential for cross-cultural historical research. Curated for fine-tuning on Ch
Smol LLMs for History
Small language models (≤3B params) that can be fine-tuned or run locally for history-focused tasks. Curated for training on Chinese and Brazilian hist
-
HuggingFaceTB/SmolLM2-135M
Text Generation • 0.1B • Updated • 2.46M • 228 -
HuggingFaceTB/SmolLM2-135M-Instruct
Text Generation • 0.1B • Updated • 1.41M • 405 -
HuggingFaceTB/SmolLM2-360M
Text Generation • 0.4B • Updated • 361k • 124 -
HuggingFaceTB/SmolLM2-360M-Instruct
Text Generation • 0.4B • Updated • 286k • 213