Gyuho Lee
United States
969 followers
500+ connections
View mutual connections with Gyuho
Gyuho can introduce you to 10+ people at NVIDIA
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Gyuho
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
Distributed systems(consensus, etcd, Kubernetes), AI.
Activity
969 followers
-
Gyuho Lee reposted thisGyuho Lee reposted this[tvN '유 퀴즈 온 더 블럭' 젠슨 황 CEO 편, 지금 바로 확인해 보세요! 📺] 지난 수요일 방영된 '유 퀴즈 온 더 블럭' 다들 본방사수 하셨나요? NVIDIA 젠슨 황 CEO가 전한 치열했던 일대기와 가슴속 깊은 이야기, 그리고 따뜻한 가족들의 모습까지 모두 담긴 감동적인 회차였는데요. 🙌 혹시 방송을 놓치셨다면, 지금 바로 하이라이트 영상으로 만나보세요! https://nvda.ws/4uCMP9y[#유퀴즈온더블럭] 인생 첫 예능 출연이 유퀴즈⁉️ 대한민국이 특별하게 느껴진다는 젠슨 황💖 | #티비빅[#유퀴즈온더블럭] 인생 첫 예능 출연이 유퀴즈⁉️ 대한민국이 특별하게 느껴진다는 젠슨 황💖 | #티비빅
-
Gyuho Lee reposted thisGyuho Lee reposted this현재 NVIDIA는 한국에 'NVIDIA AI 테크놀로지 센터 (NVAITC) 솔루션 아키텍트' 를 채용하고 있습니다. 한국의 대학 및 교육 기관, 대기업, 스타트업, 정부 기관의 연구원 분들과 함께 협업 프로젝트를 진행할 인재를 채용하고 있습니다. 아래 채용 공고를 살펴보고, 지금 바로 지원하세요. ✅ 채용 공고 보기: https://nvda.ws/4xoIDgk
-
Gyuho Lee reposted thisGyuho Lee reposted this한국에서 보낸 특별한 일주일 🇰🇷 김포공항에 도착한 순간부터 AI Ecosystem Reception까지, 이번 방문은 단순한 만남이 아닌 한국과 함께 AI의 미래를 설계하고 만들어가는 시간이었습니다. 🍗 LG Electronics, NAVER Corp, SK hynix 리더들과의 만찬 ✨ Faker와의 만남, T1 Base Camp 깜짝 방문, 그리고 Riot Games x RTX Spark 발표 💻 두 곳의 PC방 방문, RTX Spark 데모, 그리고 현장의 모든 게이머들과 함께한 점심 ⚾ 잠실야구장에서의 시구 🦞 서울대학교 Build-a-Claw에서 만난 차세대 AI 빌더들 한국은 단순한 파트너를 넘어, 앞으로의 AI 혁신을 함께 설계해 나가는 동반자입니다. 전체 이야기 보기 ➡️ https://nvda.ws/4x7Eg98
-
Gyuho Lee reposted thisGyuho Lee reposted this네이버는 NVIDIA DSX 플랫폼을 활용해 국내에 풀스택 NVIDIA AI 팩토리를 구축하고 있습니다. 젠슨 황 CEO는 이번 방한 기간 중 NAVER Corp의 이해진 의장과 회동을 가졌으며, 네이버는 GAK 세종 데이터센터를 55메가와트 규모로 확장하는 것을 시작으로 향후 기가와트 규모까지 확대해 나갈 계획입니다. 🔗 자세히 알아보기: https://nvda.ws/3QubVt8
-
Gyuho Lee reposted thisGyuho Lee reposted this📣 SK 하이닉스와 NVIDIA, 글로벌 AI 팩토리 구축을 위한 차세대 메모리 공동 개발 다년간 기술 파트너십 발표 SK 하이닉스는 NVIDIA Vera Rubin부터 Jetson Thor까지 NVIDIA 플랫폼에 탑재될 메모리를 공동 개발하는 한편, NVIDIA Omniverse 라이브러리를 활용한 반도체 팹 디지털 트윈을 고도화하고 NVIDIA CUDA-X와 PhysicsNeMo를 적용해 반도체 설계 및 제조를 가속화합니다. 블로그 읽기 👉 https://nvda.ws/43pIIT8
-
Gyuho Lee reposted thisGyuho Lee reposted this📣 AI가 실제 활용의 시대로 접어든 지금, NAVER는 이를 뒷받침할 인프라를 대규모로 구축하고 있습니다. NAVER Corp는 NVIDIA DSX 플랫폼을 활용해 GAK 세종 데이터센터를 확장할 예정이며, 55메가와트를 시작으로 기가와트 규모까지 확대해 한국과 그 너머를 아우르는 소버린 AI, 피지컬 AI 모델, 에이전트, 기업 워크로드를 지원할 계획입니다. 🔗 블로그 읽기: https://nvda.ws/43ksDOI
-
Gyuho Lee reposted thisGyuho Lee reposted this서울에서 기자들과 만난 NVIDIA CEO 젠슨 황과 SK그룹 최태원 회장은 AI 협력 확대 계획을 공유했습니다. NVIDIA와 SK hynix는 AI 인프라, 개인용 AI, 피지컬 AI를 아우르는 NVIDIA 플랫폼용 메모리를 공동 개발하기 위해 다년간의 협력을 추진할 계획입니다. 🔗 자세히 보기: https://nvda.ws/4of0CRX
-
Gyuho Lee reposted thisGyuho Lee reposted this두산 베어스와 함께한 잊지 못할 시간⚾️💚 NVIDIA 팀을 환대해 주시고, 젠슨 황 CEO의 시구 기회를 마련해 주신 두산베어스에 진심으로 감사드립니다.
-
Gyuho Lee reposted thisGyuho Lee reposted this[BAC가 서울대학교로 다시 돌아옵니다! ⚡] 지난번 서울대를 뜨겁게 달궜던 'Build-a-claw'가 여러분의 성원에 힘입어 다시 찾아갑니다! 🔥 NVIDIA DGX Spark 기반의 '나만의 AI 에이전트 빌드' 라이브 데모🤖부터, 임직원들과 최신 테크 및 커리어를 나누는 '네트워킹 허브'💬까지 알차게 준비 완료! 📅 6월 8일 10:00 - 14:00 다시 만날 서울대 학생들의 뜨거운 열기를 기대하며, 곧 공개될 현장 소식을 주목해 주세요! 🚀
-
Gyuho Lee liked thisGyuho Lee liked thisI just learned about the decomp scene, summarized thusly: 1) Take a binary from an old game 2) Ask AI to decompile it back into usable source code (keyword usable in case you've used last-gen decompilers) 3) Ask AI to modify it to your hearts content: new platforms, new capabilities, new mods, VR support 4) Recompile it, distribute it, ask users to supply a legal ROM for the game's assets 5) Play Halo CE in VR natively on Quest 3 and give yourself motion sickness Maybe step 5 was just me. We have devs out there who still think AI is a useless next-token predictor, and then we have non-developers out there porting 25 year old Xbox games to native VR headsets with full 6-dof room-scale and motion controls. As a 20-year software industry veteran, if you asked me to accomplish the same task a few years ago, I would have asked for a team of 20, a few years, and given us maybe a 25% chance of success. And they're doing similar ports for virtually every popular game ever made! My mind is blown.
-
Gyuho Lee liked thisGyuho Lee liked thisTechnical Blog📝: 안전한 AI 에이전트를 위한 신뢰 계층을 구축하려면? 🔐 100개 이상의 업계 파트너와 함께 OpenShell과 Sentry를 결합해 구축한 NVIDIA Open Agent Safety Platform가 공개됐습니다. 오픈소스 보안 런타임인 OpenShell은 AI 에이전트의 행동을 추적하고 정책을 적용하며, Sentry는 BlueField의 하드웨어 기반 추가 보안 계층으로 에이전트 활동을 지속적으로 모니터링하고 밀리초 단위의 격리와 차단을 지원합니다. 안전한 에이전트 시스템을 위한 개방형 생태계를 만들어 나가는 NVIDIA의 혁신을 블로그에서 자세히 알아보세요! 👉 https://nvda.ws/4yUOKZu
-
Gyuho Lee liked thisGyuho Lee liked thisWant to build AI factories with confidence? The NVIDIA DSX AI factory platform unifies AI factory design and operations across compute, networking, power, cooling, facilities, and software. 📢 We just launched DSX Ready: A new trusted program for AI factory builders, offering products and solutions that meet category-specific NVIDIA DSX requirements. The program launches with two initial categories: battery energy storage systems (BESS) and cooling distribution units (CDUs). Read more at https://lnkd.in/gE3shMG6 At launch, DSX Ready includes qualified BESS solutions Hitachi Energy, LG Energy Solution, and Tesla, and qualified CDU solutions from LG Electronics, LiquidStack and Vertiv. The landing page at https://lnkd.in/gU8KUa6r provides full details about the categories and products.NVIDIA Launches DSX Ready to Qualify Power and Cooling Products for AI FactoriesNVIDIA Launches DSX Ready to Qualify Power and Cooling Products for AI Factories
-
Gyuho Lee liked thisGyuho Lee liked this오픈소스 빌더 커넥트, 성황리에 종료 🌐👨🏻💻🔗👩🏻💻 강남의 한 독립서점에서 Avalanche Team1 코리아가 주최한 첫 번째 개발자 네트워킹 밋업 'Builder Connect'에 30여 분의 빌더가 모였습니다. 오픈소스를 주제로 모인 이번 행사에서는 블록체인과 아발란체가 무엇인지 되짚어보는 교육 세션과 함께, Rust 개발자가 이더리움 재단에 기여하게 된 여정을 소개한 스토리텔링 세션을 진행했습니다. 특별히, 마지막 네트워킹 세션에서는 참여자 한 분 한 분이 자기를 소개하며 서로의 관심사를 이해하여 자연스러운 한국식 네트워킹을 진행할 수 있도록 여건을 마련했습니다. 참여해주신 모든 분들께 감사드리며, 다음 빌더 커넥트 행사에서는 빌더들이 더욱 끈끈하게 교류할 수 있는 내용으로 찾아오겠습니다🤗 ---- 팀1 코리아는 늘 의미 있는 활동으로 건강한 웹3 커뮤니티를 만들어나가기 위해 노력하고 있습니다. 밋업에 연사로 참여해주셔서 감사합니다. [키노트] Eric의 오픈소스 개발 탐험기 Youngwoo Yang, Team1 DevRel
-
Gyuho Lee liked thisGyuho Lee liked thisIt's been two weeks since I joined Google, to bring frontier AI services to Google's dedicated and sovereign cloud. This move comes after a memorable decade at AWS, where I left in early August. I got to build Amazon EKS and later led a broader protfolio spanning EKS, ECS, ECS, Batch, and Beanstalk. I had a great team and was fortunate to have managers and mentors who bet on me early, pushed me to rise to the challenges, and gave me the room to build things that mattered to customers. That decade shaped how I think about scale, trust, and what it takes to earn customer confidence. Now I'm excited to help our most security and regulation conscious customers access frontier AI services on their terms. Early days, but the mission is clear and the team is fantastic.
-
Gyuho Lee liked thisGyuho Lee liked thisA new asset class is being born. AI factories are becoming investable infrastructure. The capital markets are mobilizing to build the infrastructure of intelligence.NVIDIA AI Factory Compute Is Becoming an Investable Asset ClassNVIDIA AI Factory Compute Is Becoming an Investable Asset ClassJensen Huang
-
Gyuho Lee liked thisRecording of my talk with Giulio Calzolari about running Slurm on Kubernetes is available on YouTube now.Gyuho Lee liked thisAs AI workloads grow, organizations often need scheduling capabilities that Kubernetes doesn't provide out of the box, such as gang scheduling, fair-share policies, and topology-aware placement. In this session, the speakers show how Slurm and Kubernetes can work together using Slinky, an open-source operator that combines proven HPC scheduling with cloud-native operations. 💡 Learn how the team operates 8,000+ GPUs for distributed AI training, what challenges they faced, and how you can bring the same approach to your own Kubernetes environment. 🎬 Watch the session here: https://lnkd.in/dAp9mP6e #CloudNative #AI #Kubernetes #opensource #community #CNSMunich
-
Gyuho Lee reacted on thisGyuho Lee reacted on thisI'm starting a new company, Intent Lab, with Junjie Bai, Xiang L., and Casber Wang, people I've worked alongside for years. We're building an autonomous team – we call it a fleet - that builds and operates software systems from a rough human intent: what you want the system to do. The fleet owns the work end to end through design, implementation, and verification, and then keeps it evolving in production. We're sharing some early results today. The fleet re-engineered the TRT-LLM inference engine and made it the fastest engine for the GLM 5.2 model. It modernized a database from a one-line prompt, building and verifying designs to pass all 6 million acceptance tests. And it's begun building new software for a world where agents use infra in completely new ways. All three came out of one general system rather than three separate efforts, which is the part that matters to us. Read our blog here: https://lnkd.in/gR8zGvnG My career started with writing software by hand: careful architecture, one system at a time. Two observations have stayed with me over the past year. Models now write code very fast. And there's still real distance between code that runs and software you'd be comfortable putting in production and owning for years. I don't think that distance is a limit of the models. It looks to me like a missing layer. Coding agents handle the writing well. What sits above them is still open: turning a vague intent into a real design, checking that the result holds up, and keeping the system evolving once it's running. That's the layer we're working on. If it works, the economics of computer engineering will change forever. For fifty years the sensible move was to build one software and sell it to as many buyers as possible, even though every buyer's workflow and constraints were different. Customers always wanted something fitted to them, and it was never possible at scale. I think the world will end up with a great deal more software, each piece shaped around someone specific. We built many open source projects like Caffe, PyTorch, ONNX, etcd, Kubernetes, and our last company Lepton AI, one system at a time. What interests me now is whether we can build the oracle that produces a thousand of such systems. If there's a system you want built, modernized, or made faster, especially one that's been sitting on the list because it's too big to staff, send me a note for early access. If you're interested in our mission, we're hiring. Let us know!
-
Gyuho Lee liked thisGyuho Lee liked thisAttackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community. During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion. That’s why we created the Open Secure AI Alliance. https://lnkd.in/g-sP-xb7Industry Leaders Unite in Open Secure AI Alliance for AI Safety and SecurityIndustry Leaders Unite in Open Secure AI Alliance for AI Safety and Security
Experience
Education
Honors & Awards
-
Volunteer of the Year 2008
Habitat for Humanity of Greater Los Angeles
Over 1,000 hours of volunteer work in Home Improvement Store at Gardena, California. Awarded Volunteer of the Year (2008), with President's Volunteer Service Award Silver Medal (2008), Congressional Award Silver Medal (2010)
View Gyuho’s full profile
-
See who you know in common
-
Get introduced
-
Contact Gyuho directly
Other similar profiles
Explore more posts
-
Esther Kim
NVIDIA • 5K followers
“Today’s glory began in the back alleys of yesterday.” Jensen Huang didn’t simply ride waves of AI— he spent 30 years building the ship capable of navigating it. From the electronics markets of Yongsan in Seoul to standing at the pinnacle of the AI era, this article from Kyungjeilbo examines NVIDIA’s journey and the essence of its success through the philosophical lens of “Go (古) · Go (枯) · Go (孤) · Go (高).” This is not a story of a fleeting spark, but of a steady light that quietly illuminates the world. Explore the full story at the link below: 🔗 https://nvda.ws/48VNqLl
10
-
Sharada Yeluri
Upscale AI • 22K followers
The news about Nvidia acquiring Groq’s assets and IP triggered a wave of speculation about it's inference roadmap. Tempted to add my 2 cents as well... 😊 Groq’s LPU architecture gives up hardware generality in exchange for a compiler-scheduled, deterministic streaming machine fed by high bandwidth on-chip SRAMs. It is a tiled architecture where interleaved vertical columns of SRAM and compute operate as a unified, synchronous grid. The on-chip SRAM has ~230MB capacity and aggregate bandwidth of ~80 TB/s (compared to ~8TB/s of HBM bandwidth in GB200). At the system level, a Groq rack typically integrates 72 LPUs connected via a switchless, low-diameter fabric (dragonfly topology). But the full rack provides only ~15GB of SRAM. In contrast, an NVL72 rack provides ~900x more memory capacity! The compiler owns instruction scheduling, data movement, and resource allocation across the entire cluster a model maps to - no dynamic scheduling, no cache misses, no network congestion. By treating memory bandwidth and deterministic dataflow as first-class design constraints, this architecture can deliver stable, low-variance latency during the decode phase, as long as the memory capacity does not become a limiting factor! But memory capacity does matter when scaling to frontier models. Storing a TB-class model would require ~70 racks worth of SRAM storage, and this does not include the additional memory needed for KV cache growth at runtime. Thus, the system-level picture changes quickly at scale. Inter-rack cabling, retimers and physical footprint all compound TCO. This is why it makes little sense for Nvidia to "replace" GPUs entirely with LPUs for inference. The memory density needed for the scale is simply not there. The pragmatic, near-term option that Nvidia could embrace, which does not require new silicon, is to absorb Groq’s static scheduling philosophy directly into it's inference compiler and runtime stack. More aggressive graph-level scheduling across kernels, tighter lifetime planning for KV cache and activation buffers, and deterministic fast paths for latency-critical decode all directly attack tail latency on existing GPUs. Medium term, Groq LPUs could connect to GPUs over NVLink (either within a tray or across the scale-up fabric). LPUs could offload speculative decoding (using draft models that reside in SRAMs) and other processing during the decode phase that can benefit from deterministic execution and ultra-low latency. Longer term, Nvidia's inference platforms will increasingly be heterogeneous. With specialized CPX-class accelerators for prefill and GPUs with tightly coupled streaming tiles for decode. Integrating Groq's LPU tiles closer to the GPU, within the package, strengthens the case for advanced 3D SRAM integration - deterministic dataflow can fully exploit the proximity and bandwidth of 3D SRAM? In short: Groq IP and SW stack don’t redefine Nvidia’s inference stack - but they can supercharge it! Any thoughts? 🤔
366
18 Comments -
Kerim Kaya
Dria • 5K followers
A few months ago our team at Dria was training Mem Agent, a 4B parameter model designed to give any LLM persistent memory through an Obsidian style markdown system. We tried a dozen combinations of base models and RL algorithms before landing on GSPO, a variant of the reinforcement learning methods that are quietly reshaping how every reasoning model gets built today. That experience gave me a deep appreciation for how much the RL training stack for LLMs has evolved in just two years, and how most of the gains came from places nobody expected. The field started with PPO, an algorithm borrowed from robotics that required four separate neural networks in memory just to train one model. Then DeepSeek had a deceptively simple idea: instead of training an entire “critic” model to judge responses, just sample multiple answers to the same question and compare them against each other. That single change (called GRPO) cut memory usage roughly in half and became the backbone of nearly every reasoning model since. But here is what I find most fascinating. The biggest improvements after GRPO were not new algorithms at all. They were subtle bugs hiding inside normalizations that everyone assumed were harmless. One team found that dividing loss by sequence length was secretly rewarding longer wrong answers. Another discovered that normalizing by standard deviation was making models obsess over problems they had already solved. A third noticed that less than 0.5% of gradient updates were causing all the training instability, and blocking just those was enough to fix everything. The pattern is clear: we went from “we need bigger, more complex algorithms” to “we need to understand what our existing algorithms are actually doing at the token level.” This is exactly the philosophy we have been building on at Dria. With Agent Alpha we showed that a 3B model trained with the right synthetic data and RL recipe can match GPT-4o on function calling benchmarks. With Mem Agent we used GSPO to teach a 4B model to read, write and reason over structured memory, rivaling models 50x its size. And with Kai Evolve we took the idea further: instead of a single RL loop, we run evolutionary optimization across hundreds of autonomous iterations, using verification as the fitness function, achieving a 192x speedup on NVIDIA B200 CUDA kernels across 540 generations. The common thread across all of this is that the bottleneck is rarely the absence of a sophisticated solution. It is almost always a subtle misunderstanding hiding inside something that already works. Alexander Weers wrote an excellent technical walkthrough of the full RL evolution from REINFORCE through GRPO and everything that followed. If you work anywhere near LLM training, it is one of the clearest overviews out there: https://lnkd.in/e_ysV-UD
11
-
Ranil Mukesh MJ
PhobosQ • 1K followers
The AI Industry Has Been Doing Reasoning Completely Wrong. I Spent 48 Hours Digging Into ByteDance’s Secret Paper… What I Found Will Completely Change How We Build AI Agents Overnight. A couple of days back I said ByteDance Seed is criminally underrated. Today I’m delivering exactly what I promised: the full SMELT teardown + code walkthrough + graphical explanation. 𝗣𝗮𝗽𝗲𝗿: SMELT -- Scaling Laws for Compute-Matched MoE Looped Transformers arXiv: 2609.01343 Authors: Shaowen Wang & team (ByteDance Seed) 𝗧𝗵𝗲 𝗰𝗼𝗿𝗲 𝗿𝗲𝘀𝘂𝗹𝘁 (𝗶𝗻 𝗽𝗹𝗮𝗶𝗻 𝗘𝗻𝗴𝗹𝗶𝘀𝗵) Can you get deeper reasoning from a looped Transformer while keeping: ● FLOPs per token identical ● Total parameters identical ● KV cache size identical Answer: 𝗬𝗲𝘀. And the advantage grows with scale. SMELT delivers: ● 18 executed layers of reasoning depth on a 12-layer physical budget ● ≤3.9% error on active FLOPs ● ≤1.0% error on stored parameters (up to 54B) ● ≤3.6% error on serving KV cache ● Up to +20.4% compute efficiency on code & complex in-context learning 𝗪𝗵𝘆 𝘁𝗵𝗶𝘀 𝗵𝗶𝘁𝘀 𝗮𝗴𝗲𝗻𝘁𝗶𝗰 𝘀𝘆𝘀𝘁𝗲𝗺𝘀 𝗵𝗮𝗿𝗱 Autonomous agents don’t just need longer context. They need 𝗿𝗲𝗰𝘂𝗿𝘀𝗶𝘃𝗲 𝗺𝘂𝗹𝘁𝗶-𝘀𝘁𝗲𝗽 𝗱𝗲𝗽𝘁𝗵 without exploding your serving bill. SMELT gives models ~50% more execution depth while keeping the exact same GPU footprint in production. That changes the economics of agent loops, tool-use chains, and long-horizon planning. I wrote the complete technical deep-dive (mechanistic breakdown + scaling surface + executable PyTorch block-recurrence code) here: 🔗 https://lnkd.in/gYJ7h4Y3 If you’re building agents, reasoning models, or serving infrastructure, this is worth 8-10 minutes. - Ranil Mukesh M J
9
-
Rene van Pelt
Bidvise • 18K followers
Anysphere is considering investment offers at a $30 billion valuation. The company, known for its coding assistant Cursor, has seen its valuation triple since mid-year, driven by fast-growing revenue. Read more: https://lnkd.in/gjKGd3EJ 📰 Subscribe to the Daily AI Brief: https://lnkd.in/eq56P3bj #aistartup #aifunding #ainews
5
-
Daniel Nkencho
Core Flux AI • 799 followers
Nvidia's iron grip just slipped. Microsoft just debuted Maia 200. And they are not playing for second place. The new benchmarks are terrifyingly good for anyone trying to compete on legacy hardware: → Outperforming Amazon's Trainium 3 → Beating Google's TPU v7 → +30% better efficiency than current standards Most business owners scroll past this. They think "This is just chips. This doesn't concern me." That is a massive mistake. This hardware war is directly driving down the cost of intelligence. It means the systems we build today will be faster, cheaper, and smarter tomorrow. Innovation is ruthless. It always prevails. The tools that seem "scary" or "too complex" right now will be the baseline requirement in six months. You can wait for the dust to settle. Or you can implement the systems that put you miles ahead while your competitors are still debating if the tech is ready. Wisdom isn't waiting. It is implementing
-
Reza Zeraat
Turing • 21K followers
Spotify's best engineers haven't written a single line of code since December. Meanwhile, AT&T's CDO just said fine-tuned small models will be the big trend of 2026. Here's what most people are getting wrong: Everyone is chasing GPT-5.3 and Claude Opus 4.6. But the companies actually winning with AI? They're fine-tuning 7B-14B parameter models that run 10x cheaper and outperform general-purpose LLMs on their specific tasks. I've seen this firsthand. I've fine-tuned domain-specific models for FinTech and HR that beat GPT-4 on task accuracy — at a fraction of the cost. The secret isn't model size. It's: → Training data quality > quantity → LoRA/QLoRA for efficient adaptation → Chain-of-thought prompting baked into fine-tuning → RAG as a complement, not a crutch → Evaluation frameworks that test real-world performance, not benchmarks 2026 is the year "smart beats big." The engineers who know how to make a small model do one thing extremely well will be worth more than the ones chasing the next frontier release. Agree or disagree? Drop your take below. #AI #LLM #FineTuning #MachineLearning #SmallLanguageModels #RAG #AIEngineering #FinTech #DeepLearning
1
1 Comment -
Simon Gillett
The Global AI Internet… • 3K followers
Most enterprises don't have an AWS problem. They have a navigation problem. 16,000+ API endpoints across AWS services. The mistake is treating that as a menu when it's actually a maze. What we see repeatedly: teams pick the wrong compute for inference, over-provision services they barely use, and burn tokens on prompts that could be 60% shorter. In our experience, the real bottleneck isn't capability — it's architectural clarity. Model selection, caching strategy, prompt engineering. We routinely cut token costs by over 50% before touching infrastructure. The unresolved question is whether most AI teams are building on AWS or just spending on it. What's the most expensive AWS architectural mistake you've encountered?
-
Aravind Raghunathan
Murugappa Group • 20K followers
India just hosted one of the world's largest AI summits, announcing a staggering ₹10 lakh crore investment in AI over the next seven years 📈. With GPU capacity expanded to 58,000 units and compute priced at just ₹65 per hour, the infrastructure is finally catching up to ambition. The event also showcased homegrown breakthroughs: Sarvam AI's Indic‑language LLMs that outperform global rivals, the launch of BharatGen – India's own multimodal model, and even PM Modi sporting Made‑in‑India AI glasses 👓. The "Create in India" mission promises to build world‑class talent and position the nation as a tech powerhouse. Yet the headlines were dominated by a controversy over a robot dog, turning a historic summit into a gossip cycle. While the world watched India pledge massive AI funding, we were busy debating a professor’s unauthorized comment. This misplaced focus reveals a deeper issue: a collective attention calibrated for scandal, not substance. The uncomfortable truth is clear – a nation gets the media it feeds. If we continue to reward noise over innovation, we will keep lagging behind. Let’s redirect our energy toward the real breakthroughs happening right now. What will you do to support India’s AI ascent? #AI #IndiaAI #TechInnovation #FutureReady #CreateInIndia #BharatGen #SarvamAI Reference: [https://lnkd.in/grE-GRun] 🔄 Share 👍 React 🌐 Visit www.aravind-r.com #AravindRaghunathan
33
3 Comments
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top contentOthers named Gyuho Lee
-
Gyuho Lee
Dallas-Fort Worth Metroplex -
Gyuho Lee
South Korea -
Gyuho Lee
Seoul, South Korea -
GYUHO LEE
South Korea -
GyuHo Lee
Seoul, South Korea
59 others named Gyuho Lee are on LinkedIn
See others named Gyuho Lee