Frank Lindseth retweeted
New tutorial | Customize your computer vision workflows with Ultralytics callbacks ⚡
Learn how to run custom logic during prediction, training, validation & export without modifying the core implementation.
Watch here ➡️ youtu.be/GW5z7HX4Jds
#MachineLearning #Python #Mlops
Frank Lindseth retweeted
HeyGen Video is live.
Our first post-trained model, on @MiniMax_AI H3. Highest quality, lowest price, starting at $0.01/sec.
We're making video creation accessible to the world.
We're releasing HeyGen Video, built for businesses that need production-quality video without production-level costs.
Pricing starts at $0.01/s through October (50% off)
Built on @Minimax_AI H3, post-trained by HeyGen.
Learn more: developers.heygen.com/heygen…
Frank Lindseth retweeted
Assess terrain from your desk with four new tools in Google Earth. goo.gle/3VWSjQF
Terrain decides where sunlight hits, how irrigation lines run, how steep an access road gets, and how much dirt has to move before anything gets planted or built. Using these tools is as easy as point-and-click:
🗺️ Elevation contours: See how the ground rises and falls with on-demand contour lines down to 1-meter intervals.
📐 Slope: Check how steep the land is at 1-meter resolution to plan access roads and building pads.
☀️ Aspect: See which way the slope faces to check sunlight exposure across your site.
⚖️ Cut and fill: Calculate how much dirt has to move so excavation matches fill and keeps hauling costs down.
These tools are available on all Google Earth plans. Standard allows you to generate one layer with a tool each month. Professional plans include expanded access and capabilities when you need more tools, a larger site, or both.
Learn more: goo.gle/4hR22AC
Frank Lindseth retweeted
GPT-6.1 Sol on Higgsfield built a prototype that gives ranchers a bird’s-eye view of their herd.
It turns drone footage into live herd counts and tracking, with alerts when an animal wanders off.
One prompt built and deployed a visual AI agent for a manufacturing line in under 30 minutes, with alerts, video search, and incident reports.
Check out our new tutorial that shows you how to build one with the new Build Vision AI skill in the NVIDIA VSS Blueprint 3.3.
Tutorial: nvda.ws/4AEyKfS
Frank Lindseth retweeted
From full 3D CT volumes to structured findings - with transparent, step-by-step reasoning.
NV-Reason-CT is an open research foundation for developers building 3D, chain-of-thought enabled medical-imaging AI.
Technical Blog: nvda.ws/4dJnvZP
Research Paper: arxiv.org/abs/2609.27511
Hugging Face: huggingface.co/nvidia/NV-Rea…
GitHub: github.com/NVIDIA-Medtech/NV…
Frank Lindseth retweeted
LLM-powered auto-annotation is now in the Ultralytics Platform! ⚡
Use models from OpenAI, Anthropic, Gemini, Kimi, Z.ai & DeepSeek to turn images into labeled training data.
Choose a model → Click Predict → Review annotations.
Try it today on the Ultralytics Platform ➡️ platform.ultralytics.com/?ut…
#Ultralytics #ComputerVision #DataAnnotation #LLM
Frank Lindseth retweeted
Object tracking explained in 3Blue1Brown style.
Learn about the tracking loop, kalman filters, why occlusion breaks and one way to fix it.
Try it out github.com/roboflow/trackers
Frank Lindseth retweeted
Build more precise vehicle detection models with oriented bounding boxes! 🚗
The Vehicles-OBB dataset uses oriented bounding boxes (OBB) to accurately annotate vehicles at various angles, making it ideal for aerial imagery, traffic monitoring, and vehicle detection at scale.
Explore the dataset ➡️ platform.ultralytics.com/mah…
#Ultralytics #YOLO26 #Vehicles
Frank Lindseth retweeted
🚨Architects are going to hate this.
Someone just open-sourced a full 3D building editor that runs entirely in your browser.
No AutoCAD. No Revit. No $5,000/year licenses.
It's called Pascal Editor.
Built with React Three Fiber and WebGPU -- meaning it renders directly on your GPU at near-native speed.
Here's what's inside this thing:
→ A full building/level/wall/zone hierarchy you can edit in real time
→ An ECS-style architecture where every object updates through GPU-powered systems
→ Zustand state management with full undo/redo built in
→ Next.js frontend so it deploys as a web app, not a desktop install
→ Dirty node tracking -- only re-renders what changed, not the whole scene
Here's the wildest part:
You can stack, explode, or solo individual building levels. Select a zone, drag a wall, reshape a slab -- all in 3D, all in the browser.
Architecture firms pay $50K+ per seat for BIM software that does this workflow.
This is free.
100% Open Source.
Frank Lindseth retweeted
The launches continue! Level up your learning with Interactive Learning Overviews— now available to ALL users! 📓✨
Under Reports, seamlessly combine source summaries with studio artifacts into one interactive hub. Perfect for studying or one stop deep dives into a new topic.
Frank Lindseth retweeted
🚨Architects are going to hate this.
Someone just open-sourced a full 3D building editor that runs entirely in your browser.
No AutoCAD. No Revit. No $5,000/year licenses.
It's called Pascal Editor.
Built with React Three Fiber and WebGPU -- meaning it renders directly on your GPU at near-native speed.
Here's what's inside this thing:
→ A full building/level/wall/zone hierarchy you can edit in real time
→ An ECS-style architecture where every object updates through GPU-powered systems
→ Zustand state management with full undo/redo built in
→ Next.js frontend so it deploys as a web app, not a desktop install
→ Dirty node tracking -- only re-renders what changed, not the whole scene
Here's the wildest part:
You can stack, explode, or solo individual building levels. Select a zone, drag a wall, reshape a slab -- all in 3D, all in the browser.
Architecture firms pay $50K+ per seat for BIM software that does this workflow.
This is free.
100% Open Source.
What happens when you give a world model 32 images of a place we know very well?
World Labs put Atlas to the test on Voyager.
Atlas brings text, images, video, and 3D into a shared spatial context, enabling one model to generate new views, reconstruct scenes, and simulate worlds. With Voyager, you can now see that in action in real time.
Pretrained from scratch on NVIDIA Blackwell GPUs. Take a look around 👇
From 32 input images to real-time flight through @nvidia's Voyager headquarters.
Trained on NVIDIA Blackwell GPUs, Atlas uses these images as 3D spatial context to generate new views, letting you explore with pixel-perfect camera control.
Take a look around.
Frank Lindseth retweeted
New tutorial | Build a security alarm system with Ultralytics YOLO26 🚨
Learn how to detect specified objects in video and automatically trigger email alerts using SMTP.
Watch here ➡️ youtu.be/ij9cnYM0EVE
#Ultralytics #YOLO26 #surveillance
Frank Lindseth retweeted
New tutorial | Visualize activity with Ultralytics YOLO26 heatmaps 📊
Learn how to monitor the areas where activity occurs most by combining object tracking with a heatmap solution.
Watch here ➡️ youtu.be/UU87gGC-kKg
#Ultralytics #YOLO26 #Heatmaps #ComputerVision
Frank Lindseth retweeted
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Frank Lindseth retweeted
Excited to share SceneAgent 🤖, our agentic pipeline for scene simulation from 3D captures! Made by Luke Hollis and Tianxing Fan!
We use a combination of semantic features and predictive per-Gaussian physics to transform individual complex objects or entire scenes for robotics policy training and evaluation.
Our pipeline automates comprehensive details converting input images, video, lidar, or generated 3DGS scenes like from World Labs Marble.
It automates:
0. 3DGS processing from input data at 30k and 60k training steps, with automated per-image correction/calibration
1. Infer semantic features for each gaussian
2. Segment objects and infill background
3. Bake predictive physics materials for each object (rigidity, friction, density, etc).
4. Decompose objects into individual parts
5. Articulate movable pieces and joints
6. Generate similar 3d meshes with different geometries, textures, physics properties, and articulations
Connected to autoresearch tools, we believe that this will solve an important step in the time-intensive data problem for robotics training and policy evaluation.
Paper and code coming soon, more at computationalrobotics.seas.h…
Frank Lindseth retweeted
MoGe‑3 is now available in ComfyUI.
fine-detail monocular 3D geometry from open-domain images
recovers metric point/depth/normal maps + FOV in one pass
huggingface.co/Comfy-Org/MoG…
Frank Lindseth retweeted
From 32 input images to real-time flight through @nvidia's Voyager headquarters.
Trained on NVIDIA Blackwell GPUs, Atlas uses these images as 3D spatial context to generate new views, letting you explore with pixel-perfect camera control.
Take a look around.
Frank Lindseth retweeted
Computer vision can see beyond RGB. 🌈
Ultralytics YOLO26 supports multispectral object detection workflows with multi-channel TIFF datasets; learn how this unlocks richer spectral information for Vision AI.
Read more ➡️ docs.ultralytics.com/dataset…
#Ultralytics #YOLO26 #MultiSpectral #VisionAI