SenseNova U1 Pro
The flagship native multi-modal large model released by SenseTime, focusing on delivery-level complex graphic and text creation and native 8K ultra-clear output
In-depth Report
-
SenseNova U1 Pro is the flagship native multi-modal large model released by SenseTime Technology at the 2026 World Artificial Intelligence Conference (WAIC). It was officially unveiled in Shanghai on July 18, 2026. It focuses on "delivery-level" complex graphics and text creation, can natively output 8K ultra-high-definition images, and perform dozens of rounds of self-planning, generation, inspection, and correction around a complex goal. SenseTime positions it as a "delivery-level native multi-modal agent base for long-range tasks", clearly benchmarking against OpenAI's GPT-Image 2. Currently only open for invitation testing, the official version, API and pricing will not be made public until August 2026.
-
SenseNova U1 Pro is produced by SenseTime, which is headquartered in Shanghai, China. It is the leading manufacturer of visual AI in China and has accumulated more than ten years of experience in the field of computer vision. This model belongs to the flagship version of the "RiRiXin SenseNova U" series. The launch site was lectured by SenseTime Chairman and CEO Xu Li, and co-founder and chief scientist Dahua Lin gave a technical demonstration. SenseTime is following the hierarchical route of "open source to build reputation + closed source to build revenue". In April this year, it first open sourced the basic model SenseNova U1 (Lite version, Apache 2.0 protocol). By mid-July, the total number of stars of U1 and the supporting SenseNova-Skills code library on GitHub exceeded 8,500. In June this year, the average number of images generated per day per U1 user jumped to 107, nearly three times more than in May, indicating that the model has entered the real workflow. This time U1 Pro is a closed-source top-level version, supplemented by two product lines: U1 Standard (open source and lightweight) and U1 Fast (high concurrency and low cost).
-
U1 Pro is designed around "a complex goal, continuous understanding, planning, execution, inspection and correction", and the official summarizes four core capabilities. The first is professional-level design aesthetics. The image generation has been advanced from "lifelike details" to "composition, color, and layout have a sense of design." The goal is to weaken the "AI flavor" and directly reach the deliverable level. The second is native 8K output, which supports up to 7680×4320 resolution direct output. The pixels are 4 times that of 4K. Text, lines, and icons are still clear after enlargement. It is said to be able to withstand printing and exhibition-level inspections. The third is high-density graphics and text control, which can still maintain overall layout in a layout with a large amount of information, and the text rendering error rate is very low. The fourth is long-term Agentic closed-loop thinking. It can run dozens of rounds of "generation loops" around complex goals. It can self-check while drawing. If you find that icons are misplaced or text is misspelled, you can correct it yourself. It can also make simultaneous and precise edits to the overall style and local text. In actual demonstrations, U1 Pro has delivered three hard-core cases: the 4:1 ultra-wide format ink scroll "Intellectual World Map" drawn for the ninth anniversary of WAIC, which strung the key nodes of the nine conferences from 2018 to 2026 into an oriental aesthetic map; the one-time output of the complete world view setting of "Shadow Blade" plus 22 A logically coherent movie storyboard; it also cooperated with the office assistant "Little Raccoon" to predict the dark horse Cape Verde to qualify for the World Cup one month before the start, ranking second among the 12 large models participating in the prediction, and automatically generated a final data visualization report. These capabilities have been implemented in SenseTime’s office tool “Little Raccoon” and video creation platform “Seko”.
-
As of July 2026, no pricing has been announced for the U1 Pro. SenseTime made it clear that the official version will be open to the public in August 2026, when API services and commercial pricing will be simultaneously launched. Currently, the only way to get in touch is to apply for a test invitation through the SenseNova platform. There is no public API and no downloadable weights. From the perspective of business logic, SenseTime’s approach is to monetize in layers: U1 Pro is closed source and provides top-level 8K delivery-level capabilities to earn revenue; U1 Standard and Lite are open source and rely on community co-construction to build reputation; U1 Fast has high concurrency and low cost, and is suitable for large-scale calls. The basic SenseNova platform has previously launched a Token Plan, with 1,500 free calls refreshed every 5 hours for each model during the first month's trial period. Target customers cover office, e-commerce, educational publishing, film and television animation, urban planning and other scenarios that require professional visual delivery.
-
Since U1 Pro is still in the beta testing stage, not many users have actually used it. Public evaluations mainly come from the media and actual testing by a small number of experiencers. The overall reputation is positive. What has been repeatedly praised is its stability in ultra-long images and high information density layouts. For the same "Twenty-Four Solar Terms" Chinese painting long scroll prompts, U1 Pro can fully display the 24 solar terms and mark the date serial numbers one by one, while the comparison GPT-Image 2 has missing content. When making academic posters, the layout information output by U1 Pro with one click is extremely dense, and it also has a recognizable QR code, which is recognized for its commercial completion. Some testers also mentioned that the advantage of the model is not only the 8K resolution, but also the ability to "organize the picture around a complete goal", which can put information, layout, characters, materials and styles into the same task. However, because the public has not yet used it, the sample size of these evaluations is limited, and the stability under real workflows needs to be verified on a larger scale.
-
The media generally regards U1 Pro as a major entrant in the “design” track. SenseTime’s own judgment is that after programming capabilities, design is becoming the next main battlefield for top multi-modal models, covering scenarios such as product development, graphic design, industrial design, video production, and urban planning. At the technical level, the focus of the industry is the NEO-Unify architecture. This architecture was developed by SenseTime and Nanyang Technological University. Its biggest feature is "subtraction" - it removes the visual encoder and VAE (variational autoencoder) in the traditional multi-modal model, allowing pixels and text to enter the same Mixture-of-Transformer (MoT) backbone network. Understanding and generating share a set of representations, and there is no more "translation seam". The reason why dense small characters are easily blurred in traditional models is precisely because of these two seams, but NEO-Unify deletes them. Lin Dahua described this unification as "thinking, understanding and creation are unified in one brain, just like the screenwriter and director integrated into one".
-
The main question is: almost all comparison data currently come from SenseTime’s own demonstrations, and there is no third-party independent benchmark test. U1 Pro's claims against GPT-Image 2, the resolution advantage of 8K over 4K, and the "extremely low" text rendering error rate are all still stated by the manufacturer and have not yet been externally verified. The parameters of the model, training details, rate limits, and production environment delays have not been announced. There is not even a technical report exclusive to U1 Pro. Only papers on the underlying NEO-Unify architecture have been published. Second is the availability threshold. The number of places in the invitational testing phase is limited, and ordinary users will not be able to get it in the short term. The official pricing will have to wait until August, and the price/performance ratio is still unknown. In terms of functionality, U1 Pro currently cannot directly generate 3D model files and focuses on ultra-high-definition 2D assets. Scenes that require 3D still require Blender or a specialized 3D generation model. In addition, SenseTime positions the model as "Understanding·Generation·Action", but the current demonstration is basically focused on static image generation, and the so-called "action" capabilities (tool calling, Agentic workflow) have not been truly demonstrated. It is worth watching whether this story can be realized.
-
U1 Pro is suitable for professional users who have high requirements for graphic delivery quality: professionals and designers who need to make infographics, business PPTs, promotional posters, academic posters, and popular science illustrations; creative practitioners who want to do film and television storyboards and character concept design; and teams in e-commerce, educational publishing, etc. that require batches of high-quality visual output. Its core value is "less card draw" - if you can really get the text, layout, and composition right the first or second time in 8K, you can save the time and cost of repeatedly adjusting the image. The unsuitable scenarios are also clear: those who need to output 3D models, those who pursue instant, low-cost and fast drawings (this is more suitable for U1 Fast), or those who want stable API access to production now, it is best to wait until the official version and pricing are released in August before evaluating. As for alternatives, you can look at GPT-Image 2, Seedream, Nano Banana Pro, etc. overseas. Domestic products such as Jimeng are also on the same track.
-
The ambition of SenseNova U1 Pro is not to "draw better-looking", but to push AI drawing from "auxiliary toys" to "design productivity agents" - using native unified architecture, 8K ultra-clear, long-term self-review and high-density graphics and text control to gnaw on the hard nut of "delivery grade". The technical route is imaginative, and the demonstration effect is eye-catching, but the quality ultimately depends on the August pricing, API performance, and whether third-party independent evaluation can confirm the advantages of those manufacturers. Until then, it's more like a convincing "trailer" than a finished product that can be immediately put on the production line.
User Reviews
-
BryanPhillips_X—After getting the invitation test qualification and trying it for a round, the most intuitive feeling is that there is no need to draw cards. In the past, I used other models to produce posters. A dozen of usable pictures were discarded, and when I zoomed in, the text was completely blurred. U1 Pro ran five sets of brand posters. Four sets of compositions and colors were delivered directly. Only one set had a conservative layout, but it did not require major changes. For the first time in half a year as a designer, I didn’t edit the layout of AI-generated images in Photoshop. -
Carolyn.Coleman_77—8K native output is really good at this point, not the kind of super-resolution amplification in the later stage. I made a 4:1 ultra-wide long information image with more than 200 Chinese characters and dozens of icon modules crammed into it. When you zoom in to some parts of the image, the text lines are still clear and stable, and they are not blurred at all. Those who make exhibition boards, outdoor large screens, or printed materials no longer have to choose between clarity and information density. -
马欣洁—After all, it’s still an invitation-only test. Ordinary people can’t touch it now. Let’s wait until August. -
MoonWalkerAllen—The pricing has not been announced yet, but it is said that it will be released together with the API in August, so it depends on the price/performance ratio. -
bigelephant720—The long ink scroll of WAIC's 9th Anniversary has taken over the screen. At first glance, I thought it was a hand-painted painting like Along the River During the Qingming Festival. The 4:1 ultra-wide frame is full of characters, landscapes and pavilions in the city. Zoom in to see the officials on horseback, the musicians playing the pipa, and the merchants selling spices. Each character has a different look, movement, and costume. There is no perfunctory feeling of copying and pasting like AI. Moreover, the timeline runs smoothly from 2018 to 2026, which is really amazing. -
KHughes52037—What struck me the most was that the combination of pictures and text was finally reliable. In the past, AI overturned as soon as it touched Chinese characters, and the title and logo of the infographic were all garbled. U1 Pro's text rendering accuracy is much higher, so I basically don't need to check it word-for-word when making report PPT pictures, but I will still scan it to be on the safe side. -
PAriv—The long picture of the 24 solar terms is indeed complete. The order and date of the 24 solar terms are correct. Likewise, the prompt word GPT-Image-2 is missing. SenseTime is really dominant in this area of infographics. -
NAmil—According to actual measurements, its advantage is not just 8K, but the ability to pack information, layout, characters, materials, and styles into the same task around a complete goal. This overall organizational ability is much more difficult than simply stacking pixels. -
AaronWrightQ—Objectively speaking, it is not without shortcomings. Particularly complex cross-modal reasoning occasionally suffers from local logic breaks. There is still room for optimization in the word spacing of extremely long texts. The stability of Agentic loops under extremely long-range tasks also needs to be re-examined. But this is a normal boundary in the early stages of the new paradigm, not an architectural problem. -
冬雪_8—I would like to ask if 3D is available now? I think Bianjie still focuses on 2D ultra-high-definition. The game’s original painting can be storyboarded, but the 3D model must be matched with Blender, right? -
John_Ortiz_Max—The stance of benchmarking GPT-Image 2 is very clear, but the current five sets of comparisons are all SenseTime's own demonstrations. None of the third-party independent evaluations have been released yet, and the parameters have not been announced. Let's wait and see. -
JaimeSantana—I can make a draft by myself, review it and then refine it. This "long-term self-review" is much more reliable than producing a picture at once. -
Julie.Butler01—NEO-unify's architectural idea is quite ruthless. It directly cuts off the visual encoder and VAE, allowing pixels and text to be modeled natively on the same MoT backbone. In the past, multi-modal model understanding and generation were two systems, barely connected by the alignment module in the middle. Dense small words were easily blurred because they were stuck at the seams of these two translations. Now sharing a set of representations, native unification has indeed solved many old problems from the root, and text rendering is much more stable. This step is more interesting than heaping parameters. -
Mason.Chavez369—The strategy of open source to build reputation and closed source to generate revenue is quite smart. U1 Lite was first launched on GitHub and already has more than 8,000 stars. Pro is closed source and collects the highest amount of money. The rhythm is well controlled. -
GRamirezQ—The information density of the academic poster is really high. The architecture diagram benchmark table also comes with a scannable QR code. The GPT version has too much white space and cannot be dense. Those who are engaged in scientific research and posters can do it. -
STurnerIII—I am looking forward to it as an e-commerce artist. Main poster images and event banners have always been a pain point. Dozens of them need to be produced a day during the peak season. Using other AI to produce pictures and adjust text layout requires repeated efforts. If it can really produce a commercial finished product in one version as advertised, and the text will not be garbled and the format will not be overturned, it will save a lot of time in changing the drawings. Just be afraid of queues and limited availability of the official version, and don’t let the price in August be too outrageous. -
purpleduck186—I thought the saying goodbye to the AI flavor was marketing rhetoric at first, until I saw that set of coffee brand posters, with low saturation and warm brown with a touch of orange and red, with enough white space and restrained and uniform fonts. They were tasteful, not purely realistic. -
MrCameronZhang_2024—This wave of domestic production plans is really on the rise. -
STmit—It can already be used in Little Raccoon and Seko. The two lines of office and short drama storyboarding are connected, and the landing scene is more realistic than the pure hair model. -
BrittanyKellyII—16000×24000 director-level storyboards, 40 to 60 frames with scene and camera markings. This specification sounds outrageous, and independent creators do not need to draw cards shot by shot. -
JeffreyButler—When I saw it at WAIC, Xu Li's long scroll that covered the entire reception wall was so shocking that I couldn't believe it on the spot. -
Cheryl_AdamsK—Being able to interact does not mean being able to deliver. This sentence is indeed accurate, and it hits the biggest illusion of the biennial graph model. I just hope that not only does the press conference look good, but that the stability and price of the official version can keep up. -
Alice_SanchezIII—Lin Dahua calls it the "creation model" and says that what is delivered is a designer who lives in numbers. This analogy is quite accurate. It is the same logic as the coding model of delivering a programmer who can type code.