• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Sunday, October 4, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

New AI Transformer Learns to See Interior Design Styles the Way Humans Do

Bioengineer by Bioengineer
October 4, 2026
in Technology
Reading Time: 6 mins read
0
New AI Transformer Learns to See Interior Design Styles the Way Humans Do
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

A team of computer scientists has built an artificial intelligence system that can look at a photograph of a living room or bedroom and correctly identify its interior design style, from Art Deco to Scandinavian, with an accuracy that substantially outperforms existing models. The system, called SAT³ (Style-Aware Attention Transformer), was developed by researchers at RMIT Vietnam and RMIT Melbourne and described in a study published in Discover Artificial Intelligence. What makes the achievement notable is not simply the raw accuracy figure, but the way the model arrives at its answers: instead of fixating on obvious objects like sofas and tables, it has been engineered to notice the subtle things that actually define a style, such as wood grain, fabric weaves, color palettes, and the rhythm of decorative motifs.

The problem the researchers set out to solve is deceptively hard. Human observers can usually tell a Scandinavian interior from a Minimalist one, but the differences live in fine details: the warmth of the wood, the texture of the textiles, the restraint of the ornamentation. Conventional image-recognition systems struggle here because they were built to identify objects, not aesthetics. Convolutional neural networks, the workhorses of computer vision, are excellent at picking up local textures but lack the ability to reason about how an entire room is composed. Vision Transformers, a newer architecture that processes an image as a sequence of patches and lets every patch attend to every other, capture global spatial relationships well but tend to overlook the fine-grained texture information that styles depend on. Worse, both families of models tend to latch onto dominant semantic objects, so a model may conclude a room is Industrial simply because it sees a metal lamp, even when the overall aesthetic is something else entirely.

SAT³ addresses this by combining the strengths of both architectures and then adding something neither has: an explicit awareness of style. The pipeline begins with a convolutional stem that extracts low-level texture maps and two complementary style descriptors. The first is a texture descriptor computed through second-order pooling, which captures the co-occurrence statistics of local features, effectively encoding how materials repeat and vary across the image. The second is a palette descriptor, a normalized histogram of the image’s colors in CIELAB space, a color model designed to match human perception. Together these descriptors give the network a statistical summary of what the room looks and feels like before the heavy reasoning begins.

Those descriptors then feed into the heart of the architecture, the Style Attention Module, or SAM. In a standard Vision Transformer, attention scores are computed purely from the content of the image patches. SAM modifies this computation in two ways. First, the texture and palette descriptors are projected into the attention space and passed through a gating function, producing a vector that channel-wise modulates the queries and keys, amplifying the feature dimensions most relevant to style. Second, a learned style-affinity matrix is added directly to the attention logits, biasing the model to let patches that share stylistic similarity interact more strongly. The result is an attention mechanism that asks not just what is in the image, but which parts of the image belong together aesthetically. A dedicated Style Token Fusion step then aggregates patch-level evidence under style-guided attention into a single compact embedding that drives classification.

Architecture alone was not enough, because fine-grained style data is scarce and expensive to label. The team therefore built a two-part data strategy. They curated StyleReal, a dataset of 2,627 high-quality interior images manually annotated by experts across five categories: Art Deco, Hi-Tech, Indochine, Industrial, and Scandinavian. Annotation followed a two-stage consensus process in which two experts labeled independently and a senior annotator resolved disagreements, achieving an inter-rater reliability above 0.85 on Cohen’s kappa. To expand beyond this limited pool, the researchers trained a strong baseline model on StyleReal and used it to pseudo-label web-crawled images, keeping only predictions with confidence above 0.90 and applying a margin-based filter that discards ambiguous cases where the top two predicted classes are nearly tied. Duplicate and near-duplicate images were removed using perceptual hashing before labeling. The resulting extended dataset, StyleExt, grew to 4,791 images with a more balanced distribution across categories.

The third pillar of the system is self-supervised learning. Before any style labels are used, the SAT³ encoder is pretrained on both real and pseudo-labeled images using two complementary objectives. Masked autoencoding hides roughly 75 percent of an image’s patches and asks the model to reconstruct them from the remainder, forcing it to learn the spatial and material relationships that hold a room together. Contrastive learning, following the SimCLR recipe, presents two augmented views of the same image and trains the encoder to map them to similar embeddings, building invariance to changes in lighting, viewpoint, and cropping. During training the supervised classification loss and the self-supervised loss are optimized jointly, with the self-supervised term acting as a regularizer that preserves the learned invariances. At inference time the auxiliary branch is discarded entirely, leaving a lean encoder and classification head whose computational cost remains close to that of a standard Vision Transformer.

The experimental results are striking. Across five-fold cross-validation on StyleExt, the strongest baseline, ViT-B/16, reached 76.9 percent validation accuracy and a 75.3 percent F1-score, with DeiT and the Swin Transformer close behind and CNN models such as EfficientNet-B4 and ResNet152 trailing further. SAT³, in its full configuration, achieved 83.7 percent accuracy and an 82.6 percent F1-score, gains of 6.8 and 7.3 percentage points over the best baseline. The price was modest: the model carries about 4.4 percent more parameters and roughly 12.5 percent more computation than ViT-B/16, an overhead the authors attribute to the style attention and fusion modules. Ablation studies showed that every component earns its place. Removing the Style Attention Module or the Style Token Fusion each caused marked drops in performance, while removing self-supervised pretraining was the single most damaging change, cutting accuracy to 79.8 percent.

Perhaps the most compelling evidence comes from attention heatmaps. When the researchers visualized where their model looked while making decisions, SAT³ consistently attended to textures, material surfaces, and color distributions, the walls, floors, and lighting patterns that carry stylistic identity. Baseline Vision Transformers, by contrast, concentrated on dominant furniture objects. The model also proved more robust than its competitors under Gaussian noise and under the natural occlusions that occur in real interior photography, such as furniture overlap and viewpoint truncation, suggesting it relies on distributed stylistic patterns rather than isolated semantic cues. Confusion matrices told a similar story: the full model showed the clearest separation between visually overlapping categories like Art Deco and Indochine, distinctions that challenge even trained human evaluators.

The authors are candid about the limits of their work. The pseudo-labeling pipeline depends on a fixed confidence threshold that may not suit all cultural contexts, and the comparison between domain-specific and generic self-supervised pretraining was not fully controlled for dataset scale, so the exact contribution of domain adaptation remains to be isolated. The robustness analysis used small evaluation subsets and cannot support formal statistical claims. Broader cross-domain validation on independent datasets from different cultural and architectural traditions is still needed, and the team notes that direct benchmarking against other specialized interior style frameworks remains future work as standardized benchmarks mature.

Even so, the implications reach well beyond interior decorating. The core insight, that attention can be explicitly biased toward perceptual attributes like texture periodicity and color harmony rather than object identity, applies to any visual task where style matters more than content. The researchers point to fashion categorization, artistic style identification, cultural heritage cataloging, and multimedia retrieval as natural extensions, and they envision integration with multimodal foundation models that combine visual representations with textual design knowledge. For now, the study stands as a demonstration that teaching machines to see beauty’s subtler signals is not only possible but measurable, and that the path there runs through architectures designed, quite deliberately, to care about the things that make a room feel like it belongs to a particular aesthetic tradition.

Subject of Research: Fine-grained interior design style recognition using a style-aware attention transformer

Article Title: A style aware attention transformer for fine grained interior design style recognition

Article References: Nguyen, B., Dao, S., Alavi, A., & Pham, H. (2026). A style aware attention transformer for fine grained interior design style recognition. Discover Artificial Intelligence, 6(1), Article 1293. https://doi.org/10.1007/s44163-026-02202-2

Image Credits: AI Generated

DOI: 10.1007/s44163-026-02202-2

Keywords: interior design, style recognition, attention transformer, computer vision, self-supervised learning, pseudo-labeling, fine-grained classification, Vision Transformer, masked autoencoding, contrastive learning, machine learning, aesthetics

Cite Scienmag News

APA
MLA
Chicago

Blake Davidson. (October 4, 2026). New AI Transformer Learns to See Interior Design Styles the Way Humans Do. Scienmag. https://scienmag.com/new-ai-transformer-learns-to-see-interior-design-styles-the-way-humans-do/

Blake Davidson. “New AI Transformer Learns to See Interior Design Styles the Way Humans Do.” Scienmag, 4 October 2026, https://scienmag.com/new-ai-transformer-learns-to-see-interior-design-styles-the-way-humans-do/. Accessed 4 October 2026.

Blake Davidson. “New AI Transformer Learns to See Interior Design Styles the Way Humans Do.” Scienmag. October 4, 2026. https://scienmag.com/new-ai-transformer-learns-to-see-interior-design-styles-the-way-humans-do/

Copy citation
Download RIS

Tags: advancements in style-aware AI modelsaestheticsAI transformer for interior designartificial intelligence in interior stylingattention transformercomputer visioncomputer vision for aestheticscontrastive learningdeep learning for interior design classificationdistinguishing design styles in imagesfine detail recognition in home decorfine-grained classificationhuman-like style identificationinterior designInterior design style recognitionMachine learningmasked autoencodingpseudo-labelingSAT³ style-aware modelself-supervised learningstyle detection in photographsstyle recognitionsubtle details in interior designvision transformer

Share12Tweet7Share2ShareShareShare1

Related Posts

Federated AI Learns to Spot IoT Cyberattacks Without Sharing Private Data

Federated AI Learns to Spot IoT Cyberattacks Without Sharing Private Data

October 4, 2026
Eggshells and Ceramic Trash Turned Into Ultra-Strong, Low-Carbon Cement Alternative

Eggshells and Ceramic Trash Turned Into Ultra-Strong, Low-Carbon Cement Alternative

October 4, 2026

Two Weeks of Overeating Weakens the Gut Barrier and Ignites Liver Immunity in Healthy Men

October 4, 2026

Perovskite Electrodes Could Supercharge the Next Generation of Lithium-Ion Batteries

October 4, 2026

POPULAR NEWS

  • Federated AI Learns to Spot IoT Cyberattacks Without Sharing Private Data

    29 shares
    Share 12 Tweet 7
  • Ancient Viral Fossils Awaken in DNMT3A-Mutant Blood Clones, Fueling Inflammation

    29 shares
    Share 12 Tweet 7
  • Eggshells and Ceramic Trash Turned Into Ultra-Strong, Low-Carbon Cement Alternative

    29 shares
    Share 12 Tweet 7
  • Laser Therapy Offers New Hope for Rare Spinal Tumors When Surgery Runs Out of Options

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Federated AI Learns to Spot IoT Cyberattacks Without Sharing Private Data

Ancient Viral Fossils Awaken in DNMT3A-Mutant Blood Clones, Fueling Inflammation

Eggshells and Ceramic Trash Turned Into Ultra-Strong, Low-Carbon Cement Alternative

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.