AI & ML interests

None defined yet.

Nymboย 
posted an update 1 day ago
view post
Post
1619
Introducing Inflect-v2, two exceptionally small, open-weight English TTS models at just 3.9M and 9.3M parameters. Both generate speech multiple times faster than real-time on CPU. Despite their size, Inflect-v2 delivers quality that is competitive with much larger lightweight TTS systems, including KittenTTS, Piper, and Supertonic-3.

CPU, CUDA, PyTorch, and ONNX are supported. Apache 2.0.

See it for yourselves:
owensong/Inflect-Micro-v2
owensong/Inflect-Nano-v2

Try the Demos:
Nymbo/Inflect-TTS (unlimited CPU usage)
owensong/Inflect-v2 (ultra-fast ZeroGPU usage)
  • 4 replies
ยท
Shrijanagainย 
posted an update 26 days ago
view post
Post
195
Welcome Researcher and Developers!

SKT AI Labs, we are pushing the boundaries of AI architecture and researchโ€”and today, we are thrilled to open our doors to the global research community!

โ€‹We warmly welcome researchers, developers, and AI enthusiasts to join us and contribute to our R&D efforts.

โ€‹๐Ÿงช What You Can Explore:

We invite you to experiment with our WMF (Weight Manifold Fusion) technology. You can test this high-dimensional fusion technique on smaller models to gain a deeper understanding of its behavior and token convergence.

---------- CHECK OUT:

SPACE : SKT-NRS/RD
EXPERIMENT : https://huggingface.co/sKT-Ai-Labs/SKT-SURYA-H
DIRECT TO MAIN DISCUSSION : SKT-NRS/RD#1

โ€‹๐Ÿค Your Feedback Shapes the Future :

โ€‹If it works: Fantastic! Share your results with us and contribute directly to the core vision of SKT AI Labs.

โ€‹If it doesn't work: No problem at all! Your critical feedback is just as valuable to us. Every experiment and anomaly helps us refine this architecture to make it more stable and robust.

โ€‹We firmly believe that true innovation stems from community collaboration and transparent testing. Let's build the future of advanced AI together. Your ideas, test results, and feedback are always welcome!

You Can Still Research and Development On WMF Only SKT-SURYA-H Model is Dismissed.

โ€‹Let's innovate and build together! ๐Ÿ’ก
Shrijanagainย 
posted an update 30 days ago
view post
Post
209
๐Ÿš€ Big News for the AI Community! ๐Ÿ”ฅ

Weโ€™re excited to release NRS_QWEN_MYTHOS_1M โ€” a powerful reasoning model built on Qwen 3.5 9B!
At SKT AI LABS, weโ€™ve supercharged this 9B model with our proprietary Neural Reasoning System (NRS) to deliver next-level performance.

๐Ÿ”ฅ Why This Model is a Game-Changer:
โœ… 100x Reasoning Capacity โ€” Exceptional deep logical thinking and complex problem-solving
โœ… 1 Million Token Context โ€” Perfect for massive codebases, long documents, and multi-turn agentic workflows
โœ… Advanced Thinking Mode โ€” Native <think> tags for true step-by-step Chain-of-Thought reasoning
โœ… Tool-Use Ready โ€” Optimized for Python execution, Web Search, and self-correction
โœ… Blazing Fast โ€” Runs smoothly on consumer GPUs like RTX 3090/4090

Technical Highlights:

Base: Qwen 3.5 9B
Tuning: NRS-specific high-quality reasoning data
Context: 1M Tokens (YaRN Scaling)
License: NRS DOCS

Whether youโ€™re a developer building coding agents, a researcher working with long-context data, or someone who loves powerful reasoning โ€” this model is built for you.

๐Ÿ‘‰ Try it now on Hugging Face:
SKT-NRS/NRS_QWEN_MYTHOS_1M

Drop a comment: What will you build with it first? ๐Ÿ‘‡
#AI #OpenSource #LLM #Qwen #ReasoningModel #HuggingFace #NewModel #AICommunity
Shrijanagainย 
posted an update 2 months ago
view post
Post
2627
We are pleased to announce that the W-IMG Vision Dataset infrastructure is officially live.

The complete asset infrastructure is now accessible on Hugging Face for internal validation and architecture scaling targets.

Dataset Endpoint - sKT-Ai-Labs/W-IMG

#SovereignAI #ComputerVision #MachineLearning #OpenSource
Parveshiiiiย 
posted an update 3 months ago
view post
Post
631
๐Ÿš€ Sonic: A lightweight Python audio processing library with tempo matching, BPM detection, time-stretching, resampling & track blending โ€” now with GPU (CUDA) acceleration for 10x speed!

Perfect for quick remixes, batch edits or syncing tracks.

๐Ÿ‘‰ https://github.com/Parveshiiii/Sonic

#Python #AudioProcessing #OpenSource #PyTorch
Parveshiiiiย 
posted an update 4 months ago
view post
Post
1643
Excited to announce my latest open-source release on Hugging Face: Parveshiiii/breast-cancer-detector.

This model has been trained and validated on external datasets to support medical research workflows. It is designed to provide reproducible benchmarks and serve as a foundation for further exploration in healthcare AI.

Key highlights:
- Built for medical research and diagnostic study contexts
- Validated against external datasets for reliability
- Openly available to empower the community in building stronger, more effective solutions

This release is part of my ongoing effort to make impactful AI research accessible through **Modotte**. A detailed blog post explaining the methodology, dataset handling, and validation process will be published soon.

You can explore the model here: Parveshiiii/breast-cancer-detector

#AI #MedicalResearch #DeepLearning #Healthcare #OpenSource #HuggingFace

Shrijanagainย 
posted an update 4 months ago
view post
Post
4320
sKT-Ai-Labs


Join fast we will soon published tokens and all join and get started because we will soon off join request button if you want you can join fast guys
  • 1 reply
ยท
Shrijanagainย 
posted an update 4 months ago
view post
Post
2686
โ€‹๐Ÿš€ Bharat AI Revolution ka Hissa Banein! ๐Ÿ‡ฎ๐Ÿ‡ณ

โ€‹Kya aap Bharat ko AI ki duniya mein ek nayi pehchan dilana chahte hain ?

SKT AI Labs sirf ek naam nahi, ek mission haiโ€”desh ko digital shakti dene ka aur "Viksit Bharat" ke sapne ko sach karne ka.

โ€‹Humse Kyun Judein?

โ€‹1. Desh ka Apna AI: Hum aise models bana rahe hain jo khas taur par Bharat ki zarooraton aur bhashaon ke liye hain.

โ€‹2. Open Collaboration: Hamare Hugging Face repository par hamare kaam ko dekhein, test karein aur apna yogdan dein.

3. Technological Growth: Agar aap student hain, developer hain ya tech enthusiast hain, toh hamare saath naya seekhne aur grow karne ka yeh behtareen mauka hai.

โ€‹Join here

sKT-Ai-Labs

๐Ÿ”—
sKT-Ai-Labs


โ€‹Aaiye, saath milkar Bharat AI Revolution ko aage badhate hain! ๐Ÿ’ป๐Ÿ”ฅ

โ€‹#SKTAILabs #DigitalIndia #AIRevolution #ViksitBharat #TechInnovation #JoinTheMission
Shrijanagainย 
posted an update 4 months ago
view post
Post
6926
SOME NEW HINDI + ENGLISH DATASETS

๐Ÿ”—
- sKT-Ai-Labs/HIN
- sKT-Ai-Labs/SKT-MIX
- sKT-Ai-Labs/ST-H

Download and Use And Train Models

You Can Alsoo Use ST-x-LIGHTING Module For Faster Training

pip install ST-x-LIGHT-V11
  • 2 replies
ยท
Parveshiiiiย 
posted an update 4 months ago
view post
Post
2968
Just did something Iโ€™ve been meaning to try for ages.

In only 3 hours, on 10 billion+ tokens, I trained a custom BPE + tiktoken-style tokenizer using my new library microtok โ€” and it hits the same token efficiency as Qwen3.

Tokenizers have always felt like black magic to me. We drop them into every LLM project, but actually training one from scratch? That always seemed way too complicated.

Turns out it doesnโ€™t have to be.

microtok makes the whole process stupidly simple โ€” literally just 3 lines of code. No heavy setup, no GPU required. I built it on top of the Hugging Face tokenizers library so it stays clean, fast, and actually understandable.

If youโ€™ve ever wanted to look under the hood and build your own optimized vocabulary instead of just copying someone elseโ€™s, this is the entry point youโ€™ve been waiting for.

I wrote up the full story, threw in a ready-to-run Colab template, and dropped the trained tokenizer on Hugging Face.

Blog โ†’ https://parveshiiii.github.io/blogs/microtok/
Trained tokenizer โ†’ https://huggingface.co/Parveshiiii/microtok
GitHub repo โ†’ https://github.com/Parveshiiii/microtok
Shrijanagainย 
posted an update 4 months ago
view post
Post
5661

โ€‹We are thrilled to announce the launch of SKT-OMNI-CORPUS-2T, a massive-scale, high-quality dataset designed to power the next generation of Foundation Models (LLMs) from scratch.
โ€‹Developed at SKT AI LABS, this corpus is not just a collection of data; itโ€™s a mission to decentralize high-grade AI training for regional languages and global knowledge.

โ€‹๐Ÿ’Ž Key Highlights:

โ€‹โ€ขโ€ข Massive Scale: Targeting a multi-terabyte architecture for 2T-level tokenization.

โ€ขโ€ข โ€‹Pure Quality: Curated from 500+ Elite Sources

โ€ขโ€ข โ€‹Structured for MoE: Perfectly sharded into 3.5GB standardized units (SKT-๐•ป series) for seamless distributed training.

โ€‹๐Ÿค Open for Collaboration!

โ€‹We are looking for AI researchers, CUDA engineers, and data scientists to join us in this journey of building Project Surya and the ST-X Series models. Whether it's optimization, custom tokenization, or architecture designโ€”letโ€™s build the future together.

โ€‹Explore the Dataset on Hugging Face:

๐Ÿ”— https://huggingface.co/datasets/Shrijanagain/SKT-OMNI-CORPUS-146T-V1

DSR -- ๐Ÿ”— https://huggingface.co/datasets/Shrijanagain/SKT-DSRx10000

โ€‹#AI #MachineLearning #OpenSource #IndicAI #SKTAILABS #LLM #BigData #HuggingFace #InnovationIndia
Nymboย 
posted an update 4 months ago
view post
Post
7738
We should really have a release date range slider on the /models page. Tired of "trending/most downloaded" being the best way to sort and still seeing models from 2023 on the first page just because they're embedded in enterprise pipelines and get downloaded repeatedly. "Recently Created/Recently Updated" don't solve the discovery problem considering the amount of noise to sift through.

Slight caveat: Trending actually does have some recency bias, but it's not strong/precise enough.
  • 3 replies
ยท
Parveshiiiiย 
posted an update 6 months ago
view post
Post
353
Introducing Seekify โ€” a truly nonโ€‘rateโ€‘limiting search library for Python

Tired of hitting rate limits when building search features? Iโ€™ve built Seekify, a lightweight Python library that lets you perform searches without the usual throttling headaches.

๐Ÿ”น Key highlights

- Simple API โ€” plug it in and start searching instantly

- No rateโ€‘limiting restrictions

- Designed for developers who need reliable search in projects, scripts, or apps

๐Ÿ“ฆ Available now on PyPI:

pip install seekify

๐Ÿ‘‰ Check out the repo: https:/github.com/Parveshiiii/Seekify
Iโ€™d love feedback, contributions, and ideas for realโ€‘world use cases. Letโ€™s make search smoother together!
Parveshiiiiย 
posted an update 6 months ago
view post
Post
1653
๐Ÿš€ Wanna train your own AI Model or Tokenizer from scratch?

Building models isnโ€™t just for big labs anymore โ€” with the right data, compute, and workflow, you can create **custom AI models** and **tokenizers** tailored to any domain. Whether itโ€™s NLP, domainโ€‘specific datasets, or experimental architectures, training from scratch gives you full control over vocabulary, embeddings, and performance.

โœจ Why train your own?
- Full control over vocabulary & tokenization
- Domainโ€‘specific optimization (medical, legal, technical, etc.)
- Better performance on niche datasets
- Freedom to experiment with architectures

โšก The best part?
- Tokenizer training (TikToken / BPE) can be done in **just 3 lines of code**.
- Model training runs smoothly on **Google Colab notebooks** โ€” no expensive hardware required.

๐Ÿ“‚ Try out my work:
- ๐Ÿ”— https://github.com/OE-Void/Tokenizer-from_scratch
- ๐Ÿ”— https://github.com/OE-Void/GPT
Parveshiiiiย 
posted an update 6 months ago
view post
Post
276
๐Ÿ“ข The Announcement
Subject: XenArcAI is now Modotte โ€“ A New Chapter Begins! ๐Ÿš€

Hello everyone,

We are thrilled to announce that XenArcAI is officially rebranding to Modotte!

Since our journey began, weโ€™ve been committed to pushing the boundaries of AI through open-source innovation, research, and high-quality datasets. As we continue to evolve, we wanted a name that better represents our vision for a modern, interconnected future in the tech space.

What is changing?

The Name: Moving forward, all our projects, models, and community interactions will happen under the Modotte banner.

The Look: Youโ€™ll see our new logo and a fresh color palette appearing across our platforms.

What is staying the same?

The Core Team: Itโ€™s still the same people behind the scenes, including our founder, Parvesh Rawal.

Our Mission: We remain dedicated to releasing state-of-the-art open-source models and datasets.

Our Continuity: All existing models, datasets, and projects will remain exactly as they areโ€”just with a new home.

This isnโ€™t just a change in appearance; itโ€™s a commitment to our next chapter of growth and discovery. We are so grateful for your ongoing support as we step into this new era.

Welcome to the future. Welcome to Modotte.

Best regards, The Modotte Team
Nymboย 
posted an update 7 months ago
view post
Post
3712
Genuine recommendation: You should really use this AutoHotKey macro. Save the file as macros.ahk and run it. Before sending a prompt to your coding agent, press Ctrl + Alt + 1 and paste your prompt to any regular chatbot. Then send the output to the agent. This is the actual, boring, real way to "10x your prompting". Use the other number keys to avoid repeating yourself over and over again. I use this macro prolly 100-200 times per day. AutoHotKey isn't as new or hype as a lot of other workflows, but there's a reason it's still widely used after 17 years. Don't overcomplicate it.

; Requires AutoHotkey v1.1+

; All macros are `Ctrl + Alt + <variable>`

^!1::
    Send, Please help me more clearly articulate what I mean with this message (write the message in a code block):
return

^!2::
    Send, Please make the following changes:
return

^!3::
    Send, It seems you got cut off by the maximum response limit. Please continue by picking up where you left off.
return


In my experience the past few months, Ctrl + Alt + 1 works best with Instruct models (non-thinking). Reasoning causes some models to ramble and miss the point. I've just been using GPT-5.x for this.
Parveshiiiiย 
posted an update 7 months ago
view post
Post
3606
Hey everyone!
Weโ€™re excited to introduce our new Telegram group: https://t.me/XenArcAI

This space is built for **model builders, tech enthusiasts, and developers** who want to learn, share, and grow together. Whether youโ€™re just starting out or already deep into AI/ML, youโ€™ll find a supportive community ready to help with knowledge, ideas, and collaboration.

๐Ÿ’ก Join us to:
- Connect with fellow developers and AI enthusiasts
- Share your projects, insights, and questions
- Learn from others and contribute to a growing knowledge base

๐Ÿ‘‰ If youโ€™re interested, hop in and be part of the conversation: https://t.me/XenArcAI
  • 12 replies
ยท
Nymboย 
posted an update 7 months ago
view post
Post
2852
๐Ÿšจ New tool for the Nymbo/Tools MCP server: The new Agent_Skills tool provides full support for Agent Skills (Claude Skills but open-source).

How it works: The tool exposes the standard discover/info/resources/validate actions. Skills live in /Skills under the same File_System root, and any bundled scripts run through Shell_Command, no new infrastructure required.

Agent_Skills(action="discover")  # List all available skills
Agent_Skills(action="info", skill_name="music-downloader")  # Full SKILL.md
Agent_Skills(action="resources", skill_name="music-downloader")  # Scripts, refs, assets


I've included a music-downloader skill as a working demo, it wraps yt-dlp for YouTube/SoundCloud audio extraction.

Caveat: On HF Spaces, Shell_Command works for most tasks, but some operations (like YouTube downloads) are restricted due to the container environment. For full functionality, run the server locally on your machine.

Try it out ~ https://www.nymbo.net/nymbot
Nymboย 
posted an update 8 months ago
view post
Post
5308
๐Ÿš€ I've just shipped a major update to the Nymbo/Tools MCP server: the Agent_Terminal, a single "master tool" that cuts token usage by over 90%!

Anthropic found 98.7% context savings using code execution with MCP, Cloudflare published similar findings. This is my open-source implementation of the same idea.

# The Problem

Traditional MCP exposes every tool definition directly to the model. With 12 tools, that's thousands of tokens consumed *before the conversation even starts*. Each tool call also passes intermediate results through the context window โ€” a 10,000-row spreadsheet? That's all going into context just to sum a column.

# The Solution: One Tool to Rule Them All

Agent_Terminal wraps all 12 tools (Web_Search, Web_Fetch, File_System, Generate_Image, Generate_Speech, Generate_Video, Deep_Research, Memory_Manager, Obsidian_Vault, Shell_Command, Code_Interpreter) into a single Python code execution gateway.

Instead of the model making individual tool calls, it writes Python code that orchestrates the tools directly:

# Search for Bitcoin price
result = Web_Search("current price of bitcoin", max_results=3)
print(result)


Don't know what tools are available? The agent can discover them at runtime:

print(search_tools('image'))  # Find tools by keyword
print(usage('Generate_Image'))  # Get full docs for a specific tool


The individual direct tool calls are all still there, but they can be disabled if using the Agent_Terminal. Try it now - https://www.nymbo.net/nymbot
  • 1 reply
ยท
Parveshiiiiย 
posted an update 8 months ago
view post
Post
1680
Another banger from XenArcAI! ๐Ÿ”ฅ

Weโ€™re thrilled to unveil three powerful new releases that push the boundaries of AI research and development:

๐Ÿ”— https://huggingface.co/XenArcAI/SparkEmbedding-300m

- A lightning-fast embedding model built for scale.
- Optimized for semantic search, clustering, and representation learning.

๐Ÿ”— https://huggingface.co/datasets/XenArcAI/CodeX-7M-Non-Thinking

- A massive dataset of 7 million code samples.
- Designed for training models on raw coding patterns without reasoning layers.

๐Ÿ”— https://huggingface.co/datasets/XenArcAI/CodeX-2M-Thinking

- A curated dataset of 2 million code samples.
- Focused on reasoning-driven coding tasks, enabling smarter AI coding assistants.

Together, these projects represent a leap forward in building smarter, faster, and more capable AI systems.

๐Ÿ’ก Innovation meets dedication.
๐ŸŒ Knowledge meets responsibility.